AI News

Wan 3.0: What Alibaba Actually Shipped — And What the Internet Made Up

The short version: Wan 3.0 exists. It is listed on Alibaba’s own Qwen Cloud platform as wan3.0-video, with published pricing and a live DashScope endpoint. It generates up to 30 seconds of video from text, image, audio and video inputs.

It is also not 4K, not open weights, and not independently benchmarked — three claims that dominate the first page of search results for its name.

That gap between what Alibaba has published and what the internet says about it is the real story here, and it is worth more to you than another feature list.


What is verifiably true

Alibaba Cloud launched Qwen Cloud for international markets on 26 May 2026, announced in Singapore by Alibaba Cloud CTO and International President Dr. Li Feifei. It is an AI-native model marketplace aggregating more than 150 model APIs — Qwen, Wan and HappyHorse, alongside third-party models including GLM, Kimi and DeepSeek. Its Chinese counterpart operates as Qianwen AI.

This matters because Qwen Cloud is where Wan 3.0 actually appears. The listing is first-party: assets served from Alibaba’s CDN, Alibaba’s aplus analytics stack, and API examples pointing at dashscope-intl.aliyuncs.com — the genuine DashScope international endpoint. The page carried a last-modified timestamp of 6 August 2026.

Here is what that listing states:

AttributeValue
Model IDwan3.0-video
InputsText, image, audio, video
OutputVideo
Max duration30 seconds
CapabilitiesReference, editing, replication, driving
Pricing — 480P$0.05 / second
Pricing — 720P$0.10 / second
Pricing — 1080P$0.20 / second
Concurrency2 concurrent requests
Async queue50 tasks
Rate limit30 RPM

Alibaba describes it as an all-in-one model with omni-modal reference, the ability to parse files, web pages and complex images, and “production-grade character consistency.”

That is the complete set of first-party technical detail currently available. There is no technical report, no model card, no architecture disclosure, no parameter count and no formal launch announcement.


The 4K claim is contradicted by Alibaba’s own pricing

Search results for Wan 3.0 are saturated with claims of native 4K output at 3840×2160. Several of the highest-ranking sites lead with cinematic 4K creation and comparison tables listing Wan 3.0 at “up to 4K UHD.”

Alibaba’s own pricing table has exactly three tiers: 480P, 720P and 1080P. There is no 4K tier, no 4K price and no 4K mention anywhere in the first-party listing.

A vendor that had shipped 4K would charge for 4K. The absence of a price tier is stronger evidence than the presence of a marketing claim.

One of those sites is worth examining closely, because it is a useful specimen. Its hero banner reads “Coming Soon.” Its comparison table marks Wan 3.0 as shipped and current. Its body copy says Wan 3.0 is now available. Those three statements cannot all be true. Its use-case imagery for Wan 3.0 is served from /wan-2.7/ asset paths — recycled Wan 2.7 material, relabelled. Its comparison table claims Wan 2.5 and Wan 2.6 have no API access, which is flatly wrong; Alibaba has documented API access for both since 2025. And its own footer states it is not affiliated with, endorsed by or sponsored by any original model provider.

It is an affiliate wrapper optimising against pre-release search demand. Its meta keywords include “wan 3.0 release date” and “wan 3.0 waitlist” — it is built to capture people looking for a model before it launches.


Is it open weights? No.

This is the question with the clearest answer and the most misinformation attached to it.

Wan 3.0 weights have not been published. Not on the Wan-AI organisation on Hugging Face, not on the Wan-Video organisation on GitHub, not on ModelScope. Independent trackers checking first-party channels through August 2026 consistently report the same thing: a hub-wide Hugging Face search for Wan 3.0 returns zero models, and the Wan-Video GitHub organisation contains four repositories — Wan2.1, Wan2.2, Wan-skills and a diffusers fork.

Several high-ranking pages assert that Apache 2.0 weights shipped in April 2026 in 1.3B and 14B variants. That claim traces to a single LinkedIn post and is not corroborated by any first-party channel. It is also internally suspicious: 1.3B plus 14B is precisely the Wan 2.1 configuration from February 2025. It reads like a recycled spec rather than a reported one. Other pages claim a 60B mixture-of-experts architecture. Both cannot be right, and neither is sourced.

The Wan open-weight record

The pattern matters more than any single release:

VersionReleasedOpen weights
Wan 2.1Feb 2025Yes — Apache 2.0
Wan 2.2Jul 2025Yes — Apache 2.0
Wan 2.5 (preview)Sep 2025No — API only
Wan 2.6Dec 2025No — API only
Wan 2.7Apr 2026No — API only
Wan 3.02026No — API only

Wan 2.2 remains the last video flagship Alibaba open-weighted, and it has held that position through more than a year of version bumps. If you need self-hosted Wan on your own hardware, Wan 2.2 is still the answer.

The nuance worth holding: Alibaba has not abandoned open weights generally. In July 2026 it released WanSong (text-to-music), Wan-Dancer-14B (music-to-dance) and Wan-Streamer (real-time streaming) under Apache 2.0. The strategy is legible — open the adjacent and ecosystem models, monetise the video flagship. Weigh any future “Wan will be open” pledge against that record.


How good is it? Nobody independently knows yet.

Wan 3.0 does not appear anywhere on the Artificial Analysis Video Arena, the main public leaderboard built from blind human comparisons. Not in the top 27. That is a meaningful absence — models with real availability and real quality get voted on quickly.

What we can see is where the rest of the field sits. Artificial Analysis text-to-video, with audio, as of early August 2026:

#ModelCreatorEloAPI $/min
1Gemini Omni FlashGoogle1,244$6.00
2MiniMax H3MiniMax1,238$7.80
3Dreamina Seedance 2.0 720pByteDance Seed1,224$9.07
4Wan2.7-260612Alibaba1,161$9.00
5HappyHorse-1.1Alibaba-ATH1,148$9.90
6HappyHorse-1.0Alibaba-ATH1,127$13.20
7Kling 3.0 1080p (Pro)KlingAI1,111$20.16
8Wan 2.7Alibaba1,110$9.00
11Veo 3.1Google1,093$24.00

Two things stand out.

Wan is not the frontier. Alibaba’s best-scoring video model sits fourth, roughly 83 Elo behind the leader. Wan is a strong, well-priced, mid-tier option — not a category winner. Anyone telling you Wan 3.0 is the best video model available is extrapolating from a model nobody has measured.

Alibaba is hedging across two video lines. HappyHorse — an Alibaba-ATH model line — ranks immediately behind Wan 2.7 and ahead of it on some tasks. More tellingly, Alibaba’s own Model Studio documentation recommends HappyHorse over Wan as the default for image-to-video and reference-to-video work, steering users to Wan only when they need custom audio upload. When a vendor’s own docs route you away from its headline brand, that is worth noting.

The pricing problem

Wan 3.0’s 1080P tier at $0.20 per second works out to $12.00 per minute. On the same basis, Wan 2.7 costs $9.00 per minute.

Wan 3.0 is roughly 33% more expensive than Wan 2.7, and more expensive than every single model ranked above Wan 2.7 on the leaderboard — including MiniMax H3 at $7.80 and Gemini Omni Flash at $6.00.

That is a genuine commercial question. You are paying a premium over Alibaba’s own measured model for an unmeasured one. The 30-second single-pass generation may well justify it for specific workflows — a full broadcast spot in one take, with no seams to stitch, has real production value. But “more expensive than the frontier, quality unknown” is the honest framing today.


The competitive set

Wan 3.0 arrived into the most crowded week AI video has had. Two direct competitors launched within days of each other.

Seedance 2.5 (ByteDance)

Previewed at the Volcano Engine FORCE conference on 23 June 2026 and released 31 July 2026. ByteDance skipped versions 2.1 through 2.4 to signal a generational jump.

  • 30-second single-pass audio-video generation, with multi-turn extension for longer sequences
  • Up to 50 reference inputs — 30 images, 10 video clips, 10 audio clips — the highest input ceiling in the category
  • Available on Jimeng (Dreamina) and Doubao Pro
  • API not yet live. ByteDance’s own launch post says access is coming soon via BytePlus ModelArk; enterprise documentation points to 7 August 2026
  • Closed weights
  • No published pricing, resolution spec or benchmark claim in the primary announcement

Its predecessor, Seedance 2.0, sits third on the leaderboard and separately received a 4K upgrade at the same event.

MiniMax H3 (Hailuo 3.0)

Previewed at WAIC 2026 on 17 July and launched 31 July 2026 — the same day as Seedance 2.5.

  • Native 2K (2560×1440), 24fps, 4–15 second clips, 32kHz stereo audio generated in the same pass
  • Omni-reference: up to 9 images, 3 video clips, 3 audio clips — 12 files total
  • #2 in text-to-video, #3 in image-to-video, #1 in video editing on Artificial Analysis
  • $0.13 per second for 2K ($7.80/min); a 768p tier at $0.09/s is in closed beta
  • Weights shipped 3 August 2026 — and this is the important part

MiniMax published H3 to Hugging Face as MiniMaxAI/MiniMax-H3: a 33B dense single-stream transformer, two task-specific checkpoints (FL2VA for frame-conditioned generation, Ref2VA for reference-driven), around 21GB each. The smallest working combination is 42.5GB, down 66% from 123.6GB at full precision. Native ComfyUI support merged the same day.

Three caveats that most coverage skipped:

  1. Local generation is natively 768px on the short edge, not 2K. The 2K output comes from a second in-context regeneration pass that remains hosted.
  2. H3-Context-IR and H3-Regenerate-2K did not ship. You get H3-Base; the full hosted system has three layers.
  3. The licence excludes the US, EU, UK and Korea from local deployment. The MiniMax H3 Community License, effective 2 August 2026, carves out those jurisdictions — citing the EU AI Act, UK and Korean rules, and ongoing US copyright litigation — alongside a $20M revenue threshold and an attribution requirement. MiniMax’s own repo documentation frames it as “not yet,” not “not ever.” The API remains globally available because MiniMax controls that infrastructure. Only the Qwen3-VL-32B text encoder ships under a genuine OSI licence (Apache 2.0).

If you are evaluating H3 for self-hosting, read the LICENSE file before you download anything. Canada is not on the exclusion list, but that is a determination to make against the actual document, not a blog post.


Head to head

Wan 3.0Seedance 2.5MiniMax H3
VendorAlibabaByteDanceMiniMax
Released2026, undated31 Jul 202631 Jul 2026
Max duration30s30s single-pass4–15s (extendable)
Max resolution1080PNot published2K (2560×1440)
Native audioYesYesYes, 32kHz stereo
Reference inputsOmni-modal, unspecified50 (30 img / 10 vid / 10 audio)12 (9 img / 3 vid / 3 audio)
API statusLive (Qwen Cloud)Pending (ModelArk)Live
Price (1080p/min)$12.00Not published$7.80 (2K)
Open weightsNoNoYes, with carve-outs
Independent benchmarkNoneNone (2.0 ranks #3)#2 T2V, #3 I2V, #1 editing

On today’s evidence, MiniMax H3 is the strongest overall proposition — highest-ranked of the three, cheapest per minute, and the only one you can actually download. Seedance 2.5 has the most ambitious input architecture and the best pedigree, but you cannot call it via API yet. Wan 3.0 has the longest confirmed duration alongside Seedance and a live endpoint today, which is not nothing — but it is the only one of the three with no independent quality signal whatsoever.


Wider context worth carrying

The Western default is disappearing. OpenAI discontinued the Sora consumer app on 26 April 2026, and sora-2 and sora-2-pro are scheduled to shut down on 24 September 2026 along with the entire Videos API. OpenAI’s deprecation notice names no recommended replacement. Anyone currently building on Sora has weeks, not months.

Google’s best video model is not the one it markets as its video model. Gemini Omni Flash leads the arena at $6.00/min. Veo 3.1 — the branded flagship — ranks eleventh at $24.00/min: four times the price for lower measured quality.

Chinese labs hold most of the top ten. On the with-audio text-to-video leaderboard, eight of the top ten positions belong to Chinese labs. This is the structural fact underneath every “which video model” question in 2026, and most comparison content has not absorbed it.


How this was verified

Worth stating plainly, because the method is the point.

Every capability and price claim above was checked against first-party sources: Alibaba’s Qwen Cloud model listing, Alibaba Cloud Model Studio documentation, the Wan-AI organisation on Hugging Face, the Wan-Video organisation on GitHub, ByteDance’s launch materials, MiniMax’s Hugging Face model card and licence, and the Artificial Analysis Video Arena.

Claims appearing only in aggregator content, affiliate sites or unattributed LinkedIn posts were excluded regardless of search ranking.

That filter removed a great deal — including the 4K specification, the April 2026 Apache 2.0 release, the 1.3B/14B parameter split and the 60B MoE architecture. All four are widely published. None survives contact with a first-party source.

This is a live problem for anyone evaluating AI models from search. Model-name domains now get registered and populated with generated specification pages before the model is announced, and those pages outrank primary documentation. The specs on them are plausible, internally consistent, confidently stated — and invented. When they disagree with each other, as they do here, at most one can be right.

Check the pricing page. Vendors are careless with marketing copy and careful with billing. A capability with no price tier attached usually does not exist.


What to watch

  1. Whether Alibaba formally announces Wan 3.0 at all. A live commercial endpoint with no launch post, no technical report and no leaderboard presence is unusual. It may indicate a staged rollout, a regional preview, or a naming decision still in flux.
  2. Whether Wan 3.0 appears on Artificial Analysis. That is the first genuine quality signal. Until then, the honest answer to “how good is it” is that nobody outside Alibaba knows.
  3. Whether Seedance 2.5’s API goes live on schedule. ByteDance’s enterprise docs point to 7 August 2026. That is the moment the three-way comparison becomes real.
  4. Whether Alibaba open-weights anything in the 3.x video line. The ecosystem models went Apache 2.0 in July. The flagship has not since Wan 2.2.
  5. Whether MiniMax’s jurisdictional carve-outs become the norm. If open-weight video releases start shipping with geographic exclusions as standard, “open weights” stops meaning what practitioners assume it means.

Bottom line

Wan 3.0 is a real, callable, competitively-featured video model with a 30-second ceiling and an unusually broad input surface. It is also the most over-described model in AI video right now — a genuine product buried under a layer of fabricated specification.

If you want maximum measured quality: Gemini Omni Flash or MiniMax H3.
If you want to self-host: MiniMax H3, licence permitting — or Wan 2.2, still.
If you want 30-second single takes today: Wan 3.0 is the only one you can actually call.
If you want the best answer: wait a fortnight. Seedance 2.5’s API lands, and Wan 3.0 either gets benchmarked or it does not.

Just do not pay for 4K. It is not on the menu.