AI News

Midjourney V1 vs. Google Veo 3: What Their 2025 Launches Actually Offered



Editor’s note (September 13, 2026): This article was substantially corrected after an evidence review. We corrected Midjourney’s launch access, separated launch facts from later product documentation, and removed unsupported claims about Veo 3’s failure rate, 4K output and internal architecture. The original publication date remains unchanged.

Midjourney V1 and Google Veo 3 entered the market with different starting points. Midjourney V1 was built around animating an image. Veo 3 was presented as a way to generate a scene and its sound from a text prompt. That distinction is more useful than declaring one product the winner.

This is a historical comparison of the products announced in May and June 2025. It uses the companies’ launch material and later documentation to correct the record. Kingy did not run a matched test of the two systems, so this article does not claim measured differences in quality, reliability or value.

The launch facts at a glance

Midjourney V1 and Google Veo 3 launch facts
Question Midjourney V1 at launch Google Veo 3 at launch
Main workflow Animate a Midjourney image or an uploaded image Generate a video scene and native audio from a text prompt in Gemini
Starting output Four five-second videos per job The Gemini launch post did not specify output count or duration
Extension Roughly four seconds per extension, up to four extensions Not specified in the Gemini launch post
Motion controls Automatic or manual animation; high-motion and low-motion settings Not detailed in the Gemini launch post
Audio The V1 launch post described image animation and made no audio-generation claim Native sound effects, background sound and dialogue were central launch features
Resolution Not specified in the V1 launch post Not specified in the Gemini launch post
Initial access Web-only at launch Gemini app for Google AI Ultra subscribers in the United States

Sources: Midjourney’s V1 announcement and Google’s Gemini app announcement.

These are launch conditions, not permanent product specifications. Both services changed after their announcements.

Midjourney V1 began with an image

Midjourney described V1 as an image-to-video workflow. A creator could make an image in Midjourney and select Animate, or upload an external image as the starting frame. Automatic animation asked the system to invent the motion prompt. Manual animation let the creator describe how the subject, camera or scene should move.

The two motion settings came with trade-offs acknowledged by Midjourney. Low motion suited slower scenes but could produce an output with little or no visible movement. High motion could create larger subject and camera movement, but Midjourney warned that it could also produce “wonky mistakes.” Those statements support a cautious description of the model’s limitations. They do not establish a numerical failure rate.

Each launch job produced four five-second videos. A selected result could be extended by roughly four seconds at a time, up to four times. Midjourney’s current documentation describes the resulting maximum as 21 seconds. Midjourney V1 announcement, current Midjourney video documentation.

This workflow suited a creator who already had a strong key image and wanted to explore several motions from it. It did not amount to a complete editing system. A usable production still required selection, sequencing, sound work and export checks outside the generator.

Veo 3 made audio part of the prompt

Google’s May 20, 2025 Gemini announcement introduced Veo 3 with native support for sound effects, background noise and dialogue. The launch page showed it as a text-prompt workflow inside the Gemini app, initially available to Google AI Ultra subscribers in the United States.

That launch evidence supports a simple claim: Veo 3 could generate video and audio together. It does not support the earlier article’s explanation that three named models always handled visuals, music and speech as a fixed internal architecture. Google’s public launch material cited here did not make that technical claim, so it has been removed.

The same limit applies to output quality. Google’s Gemini launch post did not promise 4K output. Later Google Cloud documentation for the veo-3.0-generate-001 and fast endpoints listed 720p and 1080p output, four-, six- or eight-second clips, and supported sound generation. Those cloud endpoints were released after the consumer launch and are now documented as retired in favour of Veo 3.1. They provide later product context, not proof of the exact Gemini settings available on launch day. Google Cloud Veo 3 documentation.

Access was described backwards in the earlier article

Midjourney’s launch post said V1 was web-only at launch. The earlier article instead described a Discord-first video workflow. That was incorrect for June 18, 2025.

Midjourney’s current documentation now describes both its website and Discord methods. A current reader can therefore use Discord, but that later availability should not be projected backwards onto the launch. The date attached to a product statement matters when an interface changes quickly.

Google’s initial consumer access was narrower. Veo 3 was available in the Gemini app to Google AI Ultra subscribers in the United States. Google announced Ultra at $249.99 per month in the same May 2025 post. That was a launch price and geographic condition. It should not be presented as a current universal price without checking the active plan page in the reader’s country.

Resolution and audio need product-specific labels

The original article paired “Midjourney at 720p” with “Veo at 4K” as though those were stable launch specifications. The retained primary sources do not support that comparison.

Midjourney’s current video documentation lists 480p as standard definition and 720p as high definition for eligible plans. The V1 launch announcement itself did not specify resolution. Google’s consumer launch announcement did not specify resolution either, while the later Cloud documentation for the Veo 3.0 endpoints lists 720p and 1080p. None of those sources supports the article’s unconditional 4K claim.

Audio is clearer. Veo 3’s launch announcement expressly included native audio. Midjourney’s V1 announcement described image animation and did not describe an audio-generation feature. A buyer comparing the services should still check the exact product surface and plan being offered now, because Gemini, Flow, Google Cloud and Midjourney do not necessarily expose identical controls or quotas.

There is no support for a 30–40% Veo failure rate

The earlier article said Veo 3 failed on 30–40% of generations and often returned silent videos. It did not identify a reproducible test, sample size, prompt set, product surface, date or source for that estimate. The number has been removed.

Generative-video systems can produce unusable results, but “unusable” needs a defined test. A visual defect, a blocked prompt, a service error and an output that misses the creative brief are different outcomes. Combining them into one percentage makes the result impossible to interpret.

A defensible comparison would run the same brief several times in each product and record:

  • Completed generations and service errors.
  • Outputs that meet the requested subject and camera movement.
  • Visible anatomy, continuity or object errors.
  • Usable audio, including speech accuracy and synchronization.
  • Time and generation cost to reach one accepted clip.
  • Product, plan, settings, date and number of attempts.

Without that record, the correct editorial position is uncertainty. Vendor descriptions can establish features. They cannot establish which tool will be more reliable for a particular production.

Compare the cost of an accepted clip

At launch, Midjourney said a video job cost about eight times as much as an image job and produced four five-second videos. It also said it would adjust pricing after observing demand. Current Midjourney documentation measures video work in GPU minutes and varies the charge by resolution and batch size.

Google tied initial Gemini access to the Ultra subscription. Google Cloud later introduced a separate developer product with its own model versions and pricing. A monthly consumer plan and metered cloud generation are different purchasing models, so a single price row can mislead.

For a real comparison, record total spend divided by accepted clips. Include discarded generations, extensions, audio repairs and editing time. A lower generation price can still cost more if the workflow needs many retries. A higher subscription price can be reasonable for frequent use, but only if the product consistently produces material the team can use.

Which workflow should you test first?

Start with Midjourney when the key creative asset is an existing image and the main question is how to animate it. Its launch workflow offered several short variations, motion settings and extensions. Test whether it preserves the subject and composition while adding the requested movement.

Start with Veo when sound belongs in the initial concept. Its launch advantage was the ability to request visuals, ambience, effects and dialogue together. Test speech accuracy, lip synchronization, continuity and whether the audio actually helps the scene.

For either product, begin with five to ten briefs drawn from the intended work. Use the same acceptance checklist and count every attempt. Check commercial terms, disclosure needs and the rights to any uploaded image, voice, likeness, logo or music before releasing the result. Midjourney’s current guidance makes the user responsible for having the necessary rights to uploaded images and for following its policies. Midjourney video responsibilities.

The evidence supports two distinct production paths: animate a chosen image with Midjourney, or generate a short audiovisual scene with Veo. It does not support calling either system universally more artistic, more realistic or more reliable. Those judgments require a dated, matched test against the work you need to produce.