Wan 3.0 Review: Features, Pricing & What's Confirmed

Wan 3.0 extends AI video to 30 seconds with 1080P output, native audio, and document to video. Here are the confirmed specs, pricing, prompts, and access routes.

Wan 3.0 Review: Features, Pricing & What's Confirmed
JXP TeamAugust 20, 202615 min read

This Wan 3.0 review looks past the launch headlines to a more useful question: what does Alibaba’s latest AI video model actually change, and which of its widely repeated specifications hold up against official documentation?

Wan 3.0 brings native clips of up to 30 seconds, output up to 1080P, synchronized audio-visual generation, and an unusually broad Omni Reference system that accepts text, images, video, audio, documents, spreadsheets, presentations, and web pages as creative input. Alibaba moved the model into public beta in August 2026.

Every specification in this Wan 3.0 review is drawn from Alibaba’s official API reference and product materials rather than secondhand spec lists. Where something is unconfirmed, it is labelled as such.

Try Wan 3.0 on JXP

Quick Summary

Wan 3.0 is one of the most workflow-oriented AI video releases of 2026. Its most important upgrades are not an unverified 4K claim or a benchmark score. They are 2-to-30-second generation with smart duration, broader reference inputs, native audio-visual output, adaptive aspect ratio handling, and the ability to treat documents and web content as source material.

For filmmakers and social creators, the longer generation window gives prompts room for camera movement, dialogue, action, and scene development. For marketers, educators, and product teams, document-to-video may matter more, because a presentation, PDF, spreadsheet, or webpage can enter the generation process directly instead of being manually rewritten into a conventional prompt.

The trade-offs are equally clear. Output tops out at 1080P in Alibaba’s official API reference, access remains in preview and may require approval, and no official Wan 3.0 model weights are currently published.

What Is Wan 3.0?

Wan 3.0 is Alibaba’s newest generation of the Wan AI video family, documented under the model ID wan3.0-video. It follows Wan 2.7, whose generation workflows were limited to a maximum of 15 seconds. Wan 3.0 officially extends that ceiling to 30 seconds in a single generation.

Alibaba describes the model as an all-in-one reference-based system supporting text-to-video, image-to-video with first-frame and first-and-last-frame control, and reference-driven generation. Rather than treating text, image, video, and audio as separate modes, Wan 3.0 combines multiple forms of reference material to direct characters, products, environments, voice, camera behaviour, and overall storytelling.

The official materials also widen the definition of a reference to include document formats and web content. A marketing team can start from a campaign deck. An educator can work from lesson materials. A product team can reference a spreadsheet alongside product images. A filmmaker can combine character references, environment images, motion references, and audio direction in one request.

For the full launch timeline and how the specs compare against earlier versions, see our Wan 3.0 release date guide.

Confirmed Wan 3.0 Specifications

These values come from Alibaba’s official Wan 3.0 API reference. Most third-party reviews cover only the headline figures, so the parameter-level detail below is worth reading closely — several widely repeated summaries are incomplete.

Parameter

Confirmed value

What it means

Model ID

wan3.0-video

One unified model across all generation modes

Duration

Integer from 2 to 30 seconds, default 5

Short 2-second tests are possible, not just 5-second minimums

Smart duration

Set duration to -1

The model recommends a length from your prompt and media

Duration with video input

Input + output ≤ 30 seconds total

Extending an existing clip eats into the 30-second budget

Resolution

480P, 720P, 1080P — default 1080P

Three tiers only. No native 4K option is currently documented

Aspect ratio

adaptive (default), 16:9, 4:3, 1:1, 3:4, 9:16

Six values, not three. Adaptive infers ratio from your media

Audio

Boolean, default true

Clips return with sound unless you disable it

Reference images

Up to 10 per request

Reference videos

Up to 5, totalling ≤ 15 seconds

Reference audio

Up to 5, totalling ≤ 15 seconds

Reference vs first/last frame

Mutually exclusive

You cannot combine reference assets with frame control

Document input

1 file or 1 web link, never both

Supported file formats

docx, doc, xlsx, xls, pptx, ppt, pdf, txt, key, pages, numbers, md

Twelve formats, including Apple iWork files

File limits

≤ 100MB, ≤ 50 pages

Page limit validated for paginated formats

Web link requirement

Publicly accessible, no login

Gated pages will not parse

Availability

Public beta, API preview

Access may require approval

Three of these are routinely reported incorrectly elsewhere: the 2-second lower bound, the six aspect ratio values with adaptive as default, and the mutual exclusivity between reference assets and first/last frame control. If you are planning production workflows, those three matter more than the headline 30 seconds.

What’s New in Wan 3.0

Native 30-Second Generation

The headline feature is the 30-second ceiling, double Wan 2.7’s 15 seconds.

Duration is not simply a matter of more seconds. Longer clips create harder consistency problems: a character has to stay recognizable, clothing and props should not gradually transform, camera moves need spatial logic, and dialogue, action, sound, and pacing have to stay coordinated across a much longer timeline.

Alibaba states the model is designed to maintain continuity while supporting more complex camera movement and continuous scenes. Independent verification of long-clip stability is not yet available, so treat 30-second coherence as a design goal rather than a measured result until third-party testing appears.

The smart duration mode is the underreported half of this feature. Setting duration to -1 hands length selection to the model, which reads the prompt and supplied media and recommends a clip length. For teams unsure how much time a scene needs, that is a cheaper starting point than guessing and regenerating.

Omni Reference

Traditional AI video workflows force a single starting point: a text prompt, a first-frame image, or a reference video. Wan 3.0 combines them.

A single generation can use one reference image to define a person, another to define a product, a video to suggest movement, audio to guide voice, and a prompt to explain how everything interacts. Alibaba states the model focuses on preserving character appearance, products, props, spatial relationships, voices, interfaces, and visual style across the sequence.

The practical constraint is the one most coverage omits: reference assets and first/last frame control cannot be used in the same request. If your shot depends on a specific opening frame, you give up multi-asset reference for that generation.

Document-to-Video

Document support is the most unusual capability in this Wan 3.0 review.

The model accepts one file or one web link per request across twelve formats, including Office documents, PDFs, Markdown, and Apple iWork files. Alibaba positions this as a way to turn static, text-heavy information into video.

That opens workflows conventional generators handle poorly. Instead of manually summarizing a 20-slide deck into a prompt, you supply the deck. Instead of describing chart values from memory, the spreadsheet becomes reference material.

This does not mean every complex document becomes a finished video. Dense tables, fine text, branding requirements, and factual narration still need checking. But for business users, the ability to turn a document to video directly may prove a bigger differentiator than cinematic image quality.

Native Audio-Visual Generation

Audio is a boolean parameter defaulting to true, so Wan 3.0 clips arrive with sound unless you turn it off. Alibaba’s materials describe dialogue, music, environmental sound, multilingual speech, voice consistency, and lip synchronization as part of the same generation pass.

This removes one of the more frustrating steps in AI video production: producing a strong silent clip, then reconstructing dialogue, effects, ambience, and synchronization afterward. It does not remove post-production. Speech quality, pronunciation, levels, timing, music suitability, and commercial rights all still need review.

Wan 3.0 Prompt Structures

A strong Wan 3.0 prompt works more like a short directing brief than a list of visual adjectives. Because the model generates longer clips, prompts can describe change over time rather than a single image-like moment.

A workable structure is:

Subject + action + environment + shot sequence + camera movement + visual style + lighting + dialogue and audio + reference instructions + continuity requirements.

The three examples below are built from Alibaba’s published prompt documentation structure and Wan 3.0’s confirmed capability limits. They are structural templates rather than tested outputs — independent verification of how strictly Wan 3.0 honours in-prompt timestamps is not yet available, so treat the timing markers as intent signals and check results against your own requirements.

Cinematic prompt

A young woman in a dark green raincoat walks alone through a neon-lit Tokyo alley after midnight. Begin with a wide establishing shot showing wet pavement reflecting red and blue signs. Slowly track behind her as she walks toward a small ramen shop. At 8 seconds, cut to a side medium shot as she stops beneath the awning and looks through the steamed window. At 16 seconds, push into a close-up as she smiles slightly. Natural rain ambience, distant traffic, soft footsteps, subtle indoor restaurant sounds, cinematic low-key lighting, realistic skin texture, stable face and clothing throughout the scene.

The important word is not “cinematic.” It is the timeline. The prompt states what happens first, what changes, how the camera moves, and what must stay stable.

Product video prompt

Use the reference product image as the exact product design. Create a premium vertical commercial in a bright modern kitchen. Begin with a close-up of the product on a stone countertop while morning sunlight moves across the surface. Pull back to reveal a creator entering the frame and picking up the product. Preserve the product’s shape, logo placement, proportions, and colour in every shot. Show one close-up demonstration of the main feature, then finish with a clean hero shot on the countertop. Natural room ambience, subtle product sounds, confident commercial pacing, realistic materials, no changes to packaging or branding.

This style tells the model which information from the reference must stay locked.

Document-to-video prompt

Use the attached presentation as the factual source for the video. Create a concise 30-second business explainer summarizing the three most important findings. Open with a clean animated title, then visualize the first key metric using the chart from the presentation. Transition to the second insight with a simple motion-graphic comparison. End with the recommended action from the final slide. Preserve all numbers and labels exactly as provided. Professional narration, restrained corporate motion design, clear information hierarchy, no invented statistics.

For document workflows, accuracy instructions matter most. Tell the model explicitly not to invent numbers or rewrite factual labels — and verify the output regardless.

Wan 3.0 Pricing

Wan 3.0 uses per-second output pricing in Alibaba Cloud Model Studio, billed on successfully generated duration.

Resolution

Official price

Approx. cost for 30 seconds

480P

$0.05/sec

$1.50

720P

$0.10/sec

$3.00

1080P

$0.20/sec

$6.00

Chinese-market listings give the equivalent tiers as ¥0.3, ¥0.6, and ¥1.2 per second.

Alibaba positions 480P for rapid exploration, 720P as the quality-efficiency balance, and 1080P as the high-fidelity option. The practical workflow follows directly: explore at lower resolution, reserve 1080P for prompts that already work. The 2-second minimum duration makes cheap structural tests possible — a 2-second 480P check costs a fraction of a full-length attempt.

Third-party platforms set their own credit pricing, which may differ because of subscription structure, queueing, or additional processing.

Wan 3.0 vs Wan 2.7

Area

Wan 2.7

Wan 3.0

Maximum duration

Up to 15 seconds

Up to 30 seconds

Minimum duration

Mode-dependent

2 seconds

Smart duration

No

Yes, via duration: -1

Documents and webpages

Not a headline capability

Core Omni Reference capability

Reference workflow

Separate generation modes

Unified Omni Reference

Adaptive aspect ratio

No

Yes, and it is the default

Maximum resolution

Up to 1080P in supported modes

Up to 1080P

Official downloadable weights

None listed

None listed

For most users the upgrade is about longer narratives, wider input types, and a more unified reference workflow — not resolution.

Limitations and Unconfirmed Claims

Access is gated. Wan 3.0 is an API preview and public beta product. Access can require approval, and capacity during preview is limited. Features, limits, and platform interfaces may keep changing.

No independent benchmarks exist. Alibaba has published strong demonstrations, but launch demos are not benchmark results. Wan 3.0 does not currently appear on public blind-comparison leaderboards. Evaluate character consistency, prompt adherence, physical motion, lip sync, text rendering, and long-sequence stability against your own requirements.

Cost compounds at full quality. A 30-second 1080P generation is listed at roughly $6. Reasonable for a finished result, expensive for iteration.

Document accuracy is not guaranteed. Supplying a PDF or spreadsheet does not ensure every number, chart, interface element, or spoken statement is reproduced correctly. Business and educational outputs need human review.

No native 4K is currently documented. Alibaba’s current API reference lists output up to 1080P. Pages describing Wan 3.0 as a native 4K model are not supported by that documentation. A platform may upscale a result to 4K, but an upscaled export is not native 4K generation.

Weights are not published. No official Wan 3.0 model-weight repository is currently listed by the Wan team, so Wan 3.0 should not be treated as a downloadable open-weight model. Wan 2.2 remains the most recent open general-purpose Wan release. We cover the evidence in detail in Is Wan 3.0 open source?

Who Should Use Wan 3.0

Wan 3.0 suits creators who regularly hit the ceiling of short AI video clips.

Filmmakers get room for previsualization, dialogue scenes, action sequences, trailer concepts, and continuous camera movement.

Marketing and e-commerce teams benefit from reference consistency and document input when working with products, campaign briefs, presentations, and vertical ads.

Educators and business teams have the most distinctive case, because slides, PDFs, spreadsheets, and webpages become creative references rather than material that must first be rewritten by hand.

Social creators can build complete TikTok, Shorts, and Reels concepts in one generation instead of assembling them from unrelated five-second clips.

Skip it if your main requirement is local inference or downloadable weights, if you need precisely rendered in-frame text, or if you need guaranteed capacity today — preview access is not a production SLA.

How to Use Wan 3.0

There are two routes: Alibaba’s own surfaces, or a hosted platform that calls the model for you.

Official Alibaba access. The model went into public beta on Alibaba Cloud Model Studio, also known as Bailian, where developers call it by the model ID wan3.0-video with an API key. Consumer-facing entry points include the Wanxiang official site, Wanjing Yike, the Qwen creation desktop client, IF STUDIO, and Duiyou, with mobile access through the Qwen app arriving as a gradual rollout rather than to all accounts at once. Region, account status, and invitation approval all affect whether the model appears in your console.

Hosted access. A hosted platform removes the Model Studio setup, API key configuration, and per-second billing account. You supply a prompt and any reference material through a browser and the generation runs remotely.

JXP’s hosted Wan 3.0 generator exposes the 480p, 720p, and 1080p tiers, the duration range up to 30 seconds, and the aspect ratio options in one interface — no Alibaba Cloud account and no API key required. Because the model is still in preview, capacity is limited and generations may occasionally need retrying.

Generate with Wan 3.0 on JXP

Whichever route you pick, the cheapest way to start is a short low-resolution test. The 2-second minimum duration means a structural check at 480P costs a fraction of a full-length attempt, so validate your prompt structure first and reserve 1080P for the version you intend to keep.

FAQ

Is Wan 3.0 released?

Yes. Alibaba announced the Wan 3.0 public beta in August 2026, delivered through Model Studio with API preview access.

How long can Wan 3.0 videos be?

Between 2 and 30 seconds, with a default of 5. Setting duration to -1 lets the model recommend a length. With video input, input plus output must not exceed 30 seconds total.

What resolution does Wan 3.0 support?

480P, 720P, and 1080P, with 1080P as the API default. These are the only tiers currently documented.

Does Wan 3.0 support 4K?

No. The official API reference currently lists three tiers ending at 1080P. Native 4K is not documented at this time, so claims of native 4K output are unsupported.

What aspect ratios does Wan 3.0 support?

Six values: adaptive, 16:9, 4:3, 1:1, 3:4, and 9:16. Adaptive is the default and infers a ratio from your input media and generation intent.

Does Wan 3.0 generate audio?

Yes. Audio is a boolean parameter defaulting to true, covering dialogue, voice, music, environmental sound, and synchronized character performance.

How many reference assets can Wan 3.0 use?

Up to 10 images, 5 videos totalling no more than 15 seconds, and 5 audio clips totalling no more than 15 seconds. Reference assets cannot be combined with first-frame or last-frame control in the same request.

Can Wan 3.0 turn documents into videos?

Yes. It accepts one file or one web link per request across twelve formats including docx, xlsx, pptx, pdf, txt, key, pages, numbers, and md, up to 100MB and 50 pages.

Is Wan 3.0 open source?

No official Wan 3.0 model-weight repository is currently listed. Wan 2.2 remains the clearly documented open-weight Wan generation.

How much does Wan 3.0 cost?

$0.05 per second at 480P, $0.10 at 720P, and $0.20 at 1080P — roughly $1.50, $3, or $6 for a 30-second clip.

Is Wan 3.0 free?

Wan 3.0 uses metered per-second API pricing. Alibaba’s Wan 3.0 pricing page does not currently advertise a dedicated free tier, although eligible Model Studio accounts may receive platform-level free quotas for supported models. Hosted platforms may also offer their own trial credits.

How do I get access to Wan 3.0?

Through Alibaba Cloud Model Studio with an API key, through Alibaba’s consumer surfaces such as the Wanxiang site, Wanjing Yike, the Qwen creation client, IF STUDIO, and Duiyou, or through a hosted platform that calls the model on your behalf. Access is still preview-gated, so availability varies by account and region.

Final Assessment

Wan 3.0’s most meaningful improvements are practical rather than numerical. It doubles the maximum clip length to 30 seconds while adding a 2-second floor for cheap testing, extends multimodal control through Omni Reference, generates synchronized audio and visuals by default, defaults to adaptive aspect ratio handling, and introduces document and webpage inputs that change how business, educational, and marketing videos get produced.

It also deserves more careful coverage than many early Wan 3.0 articles provide. The maximum output currently documented is 1080P, not native 4K, and Wan 3.0 is not currently an open-weight release. Independent quality benchmarks do not yet exist, so claims about where it ranks against competing models should be treated as unverified.

For creators who value longer coherent scenes, reference consistency, product control, native audio, and the ability to turn more than a text prompt into video, Wan 3.0 is one of the more consequential AI video releases of 2026. The best reason to use it is not the biggest specification number. It is that the model treats an entire collection of creative material — images, video, audio, documents, data, and webpages — as one directing brief.