Create with Wan 3.0, then compare Wan AI models for text, image, reference, audio, and video editing workflows.
Reference media*
Model
Resolution
Aspect ratio
A light, dreamy live-action commercial: a young European-American girl blows bubbles alone on a sunny beach. The visual is bright and breezy in vertical format, built from clear sea-sky blue, golden side backlight, pale beige sand, and iridescent highlights flowing across soap-bubble surfaces. A six- or seven-year-old girl stands at the edge of a windy seaside boardwalk, pale blonde hair continuously lifted by the wind, wearing a light summer dress and white sporty sandals, holding a small bubble solution bottle and a round children's wand. She looks excited and focused: first dipping carefully, then looking up and blowing gently, countless bubbles surrounding her as she laughs and reaches to pop them, continuously interacting with bubbles and sea breeze. Bubble formation is explicit: the clear bottle's solution gleams wet in the sun, the wand lifted from the bottle carries a complete iridescent film with a droplet hanging from its rim; the child exhales evenly, the film first tightens into a tiny mirror with rainbow interference patterns, then slowly bulges into a hemisphere, and finally a large main bubble detaches followed by two or three smaller ones carried away by the sea wind. Once free, bubbles are immediately pushed sideways and forward by the wind, their paths longer, driftier, and more unstable; bubbles of different sizes drift past in layered depth, some grazing the girl's cheek, others rising above her head. The 15-second mini-story is complete: opening three seconds in close-up, wind lifting her bangs and skirt as she opens the bottle, dips the wand, and lifts it, the camera clearly showing the full iridescent film. Then she raises the wand and blows the first big bubble and a few small ones; as they leave the wand the wind carries them upward, sea sparkles and distant wave lines reflected on their surfaces, and she laughs with widened surprised eyes. Mid-section she reaches for the largest bubble ahead, misses, laughs, turns to watch a lower cluster float past her waist, then half-turns to chase again. The climax from 10-13 seconds: a larger bubble is blown back near her face, she touches it, and it bursts at her fingertip into tiny water droplets and a brief flash of membrane light; she pauses, then breaks into happy laughter. The environment emphasizes the seaside wind and transparent air: sparkling sea, distant low waves, wooden boardwalk railings, wind-blown beach grass, and occasional seagulls. Sunlight creates strong iridescent highlights on the bubbles, wind continuously redirects their paths and lifts her hair, skirt, and wand, the air salty and bright with moisture. The camera language is designed for vertical 15 seconds: opening with close-ups of the film and pursed lips to grab attention, mid-section with medium and full-body follow shots of her chasing bubbles sideways, using the upper half of the vertical frame for floating bubbles and the lower half for running, inserting a brief slow-motion of the fingertip pop, and ending on an upward tilt that gathers sea breeze, bubbles, and the child's smiling face. Bright, rhythmic, and far more fluid than standing still to blow bubbles.
Demo 1 of 5 · Click the video to view its full prompt and settings.
What Is Wan AI?
Wan AI is a family of video generation models for creating clips from written prompts, still images, and supported reference media. The available generators range from direct text and image workflows to multimodal production and video editing, so the best version depends on the material already in your project.
The model number changes more than the label. Each version exposes a different mix of reference files, audio behavior, duration, resolution, and editing controls. Choosing by workflow first helps avoid uploading media that a particular version does not accept or selecting output settings that are unavailable in that mode.
Start with Wan 3.0 for the newest and longest reference-rich workflow. Use 2.7 when you need generation and editing on one page, 2.6 when one or two source videos should guide continuity, and 2.5 for straightforward prompt or single-image clips. The details below reflect the controls currently exposed by each generator.
Compare Wan AI Models
Start with the workflow you need, then compare inputs, reference support, audio behavior, and output options. The table reflects the controls currently available in each generator.
Wan AI model comparison
Compare by
Wan 3.0Latest model
Wan 2.7
Wan 2.6
Wan 2.5
Best fit
Longer multimodal productions
Short-form generation and video editing
Reference-led character and scene work
Direct prompt or single-image clips
Input methods
Text, first and optional last frames, images, video, and audio
Text, frames, reference videos and images, or an existing video
Text, one image, or one to two reference videos
Text or one source image
Reference media
Up to 10 images, 5 videos, and 5 audio files
Up to 5 reference files; up to 3 guide images for editing
Up to two reference videos
Single image in image-to-video mode
Audio
Audio references supported
Optional audio for text, image, and editing workflows
Native synchronized audio; no separate audio upload control
Native audio output; no separate audio upload control
Output range
2 to 30 seconds; 480p, 720p, or 1080p
Up to 15 seconds for text or image; up to 10 seconds for reference or editing; 720p or 1080p
5, 10, or 15 seconds for text and image; 5 or 10 seconds for reference; 720p or 1080p
5 or 10 seconds; 480p, 720p, or 1080p
Key advantage
Longest current workflow and broadest reference support
Dedicated video editing workflow
Focused dual-video reference workflow
Simple controls with the widest legacy resolution range
Swipe horizontally to compare every model.
Explore Wan AI Video Models
Each model has its own dedicated creation page, so you can open the right generator without losing the family-level comparison.
Wan 2.5 offers a straightforward text-to-video and image-to-video generator with flexible aspect ratios and three resolution levels.
Choose it for a familiar prompt or single-image workflow when reference libraries and video editing are not required.
Start from a written prompt or one still image
Choose 480p, 720p, or 1080p for the target channel
Wan AI FAQ
Answers based on the controls currently available in each generator.
What is Wan AI?
It is a family of video generation models with different controls for prompts, images, reference media, audio, duration, resolution, and editing. This page compares the currently available workflows before you open a dedicated model generator.
Which Wan model is the latest?
Wan 3.0 is the latest model shown here. Its generator supports text, image, and multimodal reference creation, output from 480p to 1080p, and clips up to 30 seconds when the selected references allow that duration.
Which Wan model should I use?
Use Wan 3.0 for longer clips and broad reference support. Choose 2.7 when video editing is part of the job, 2.6 for a focused one or two-video reference workflow, and 2.5 for direct text-to-video or single-image creation.
Can Wan AI create videos from text and images?
Yes. All four listed models provide text-to-video and image-to-video controls. Wan 3.0, 2.7, and 2.6 add different reference workflows, while 2.7 also provides a dedicated mode for editing an existing video.
What is the difference between Wan 3.0 and Wan 2.7?
Wan 3.0 supports the longest current output window and the broadest mix of image, video, and audio references. Wan 2.7 has a shorter output range but adds a dedicated video editing mode with prompt, image, and audio controls.
Can I still use older Wan models?
Yes. Wan 2.7, 2.6, and 2.5 each have a dedicated generator page. An older model can be the better choice when its simpler controls, reference format, duration, or resolution range matches the project more closely.