Released 24 August 2026 · two tiers
Wan 3.0 is Alibaba's newest video model, and audio is not a second pass — rain, room tone, and a single instrument line come out of the same render as the frames, already in sync. Every clip on this page was generated through the endpoints this generator calls.
Wan 3.0 vs Wan 3.0 Prime
Both tiers expose the same three modes, the same 2–30 second range, and the same resolutions. Prime costs about 40% more per second. The two clips below changed nothing but the tier.
Shared brief · seed 424242A street violinist plays beneath a rain-slick underpass at dusk. Passing headlights sweep across wet concrete, her bow moves in long steady strokes, ambient rain and a single violin line, slow push-in, 35mm film grain.
Same number, same words, two different stagings of the same scene: Wan 3.0 letterboxed itself into a wide crop under the overpass, Prime filled the 16:9 frame and put the traffic beside the player. Pick the tier before you go hunting for a composition — you cannot promote a Wan 3.0 clip to Prime and keep the frame.
Where the shot starts
Text-to-video is the default here. The other two both accept an image, and they are not near-neighbours: one treats your upload as frame one, the other treats it as a look to rebuild. The same still went into both cards below.

A single frame goes into both Wan 3.0 modes below. Everything you see downstream — the potter, the apron, the dusty palette — starts here.
Wan 3.0 image-to-video keeps the uploaded frame exactly and moves time forward from it: the hands keep turning the bowl, dust drifts through the window light. An optional second upload pins the last frame.
Wan 3.0 reference-to-video rebuilds a scene in the reference's likeness instead. Subject, wardrobe, and palette survive; in our clip the window wall came back as a shelf of finished pots.
Per second of finished video
Wan 3.0 bills resolution × seconds, and duration is linear: a thirty-second clip costs fifteen times a two-second one at the same resolution. The generator shows the credit estimate before it spends anything.
Wan 3.0 prompt structure
Keep the layers separate so that changing the sound does not accidentally rewrite the camera. Deep-thinking mode is off by default — turn it on when one prompt carries several subjects, a camera move, and a sound cue at once, and leave it off for a single action.
Name who or what is in frame and the one action they take. Two competing actions is where Wan 3.0 briefs start to drift.
One move: push in, orbit, handheld follow, or locked off. Add the lens character you want in the same clause.
Say what should be audible and what causes it. Unnamed audio still arrives, but it is whatever the scene suggests.
Seconds, aspect ratio, and resolution. Duration and resolution are the only two fields that move the price.
The potter keeps turning the half-thrown bowl, clay dust drifting through the window light, apron shifting slightly. Slow push-in, quiet workshop room tone, no music.
Wan 3.0 is Alibaba's newest Wan video model, released through inference APIs on 24 August 2026. It generates 2 to 30 seconds of video with synchronized audio from a text prompt, a still image, or reference media, at 480p, 720p, or 1080p.
Both tiers expose the same three modes, the same 2–30 second range, the same resolutions, and the same audio. Prime is the premium tier and costs about 40% more per second — $0.14 against $0.10 at 720p. Run the cheaper tier while you are still deciding the shot, then re-run the winner on Prime.
Wan 3.0 bills per second of finished video: $0.05 at 480p, $0.10 at 720p, and $0.20 at 1080p. Wan 3.0 Prime bills $0.068, $0.14, and $0.28 for the same resolutions. A five-second 720p clip is therefore $0.50 on Wan 3.0 and $0.70 on Prime, and the generator shows the credit estimate before you spend anything.
All six: text-to-video, image-to-video, and reference-to-video on both the Wan 3.0 and the Wan 3.0 Prime tier. This page opens on Wan 3.0 text-to-video; switch modes in the model picker without leaving the route.
Yes. Audio is on by default and comes out of the same render as the picture, so a scored or ambient bed arrives synchronized rather than added afterwards. Turn the audio switch off if you plan to lay your own track over the clip.
Two to thirty seconds, set as a whole number. Duration is the price lever: cost scales linearly with seconds, so a 30-second 1080p Wan 3.0 render costs fifteen times a 2-second one.
Deep-thinking mode lets the model reason over a complicated prompt before it starts rendering, which helps when a brief names several subjects, a camera move, and a sound cue at once. It is off by default and adds latency, so leave it off for simple single-action shots.
It rebuilds a scene in the reference's likeness rather than continuing that exact frame. In our demo it kept the potter, the apron, the bench, the bowl, and the dusty palette but replaced the window wall with a shelf of finished pots. Use image-to-video when the first frame must be preserved exactly, and reference-to-video when you want the look and the identity carried into a new setup.
No — a seed is only comparable inside one tier. We ran the same prompt at seed 424242 on both and got two different compositions of the same scene, so treat a tier switch as a fresh generation rather than a quality upgrade of the clip you already have.
Write a brief, pick a length and a resolution, and see the credit estimate before Wan 3.0 renders anything.
Start creatingWan 3.0 on Voor AI — Alibaba's newest video model in three modes and two tiers, from 2 to 30 seconds with synchronized audio.