Wan 3.0 on Voor AI — Alibaba's newest video model in three modes and two tiers, from 2 to 30 seconds with synchronized audio.

Example Video

Generator examples

Released 24 August 2026 · two tiers

Wan 3.0 writes the sound while it renders the picture

Wan 3.0 is Alibaba's newest video model, and audio is not a second pass — rain, room tone, and a single instrument line come out of the same render as the frames, already in sync. Every clip on this page was generated through the endpoints this generator calls.

Wan 3.0 · text to video5s · 720p · rain and violin scored in the same render
2–30sWan 3.0 clip length, set as whole seconds
1080pTop resolution, 480p and 720p below it
AudioScored in the same pass as the picture
3 modesText, image, and reference to video

Wan 3.0 vs Wan 3.0 Prime

One brief, one seed, two tiers

Both tiers expose the same three modes, the same 2–30 second range, and the same resolutions. Prime costs about 40% more per second. The two clips below changed nothing but the tier.

Shared brief · seed 424242A street violinist plays beneath a rain-slick underpass at dusk. Passing headlights sweep across wet concrete, her bow moves in long steady strokes, ambient rain and a single violin line, slow push-in, 35mm film grain.

Wan 3.0

$0.50 · 5s · 720p

Wan 3.0 Prime

$0.70 · 5s · 720p
A seed does not carry across tiers

Same number, same words, two different stagings of the same scene: Wan 3.0 letterboxed itself into a wide crop under the overpass, Prime filled the 16:9 frame and put the traffic beside the player. Pick the tier before you go hunting for a composition — you cannot promote a Wan 3.0 clip to Prime and keep the frame.

Where the shot starts

Wan 3.0 takes a prompt, a frame, or a set of references

Text-to-video is the default here. The other two both accept an image, and they are not near-neighbours: one treats your upload as frame one, the other treats it as a look to rebuild. The same still went into both cards below.

Source still of a potter holding a half-thrown bowl in a sunlit workshop
Source

One still, two destinations

A single frame goes into both Wan 3.0 modes below. Everything you see downstream — the potter, the apron, the dusty palette — starts here.

Image to video

The frame is a promise

Wan 3.0 image-to-video keeps the uploaded frame exactly and moves time forward from it: the hands keep turning the bowl, dust drifts through the window light. An optional second upload pins the last frame.

Reference to video

The frame is a mood board

Wan 3.0 reference-to-video rebuilds a scene in the reference's likeness instead. Subject, wardrobe, and palette survive; in our clip the window wall came back as a shelf of finished pots.

Per second of finished video

What a Wan 3.0 render costs

Wan 3.0 bills resolution × seconds, and duration is linear: a thirty-second clip costs fifteen times a two-second one at the same resolution. The generator shows the credit estimate before it spends anything.

ResolutionWan 3.0PrimeUse it for
480p$0.05$0.068Search the shot. Two seconds costs ten cents and still shows you whether the staging works.
720p$0.10$0.14Confirm the winner at full length. This is where you judge the audio Wan 3.0 wrote.
1080p$0.20$0.28Deliver. Send only the approved brief here, and only to Prime if it goes out publicly.

Wan 3.0 prompt structure

Four layers, one Wan 3.0 brief

Keep the layers separate so that changing the sound does not accidentally rewrite the camera. Deep-thinking mode is off by default — turn it on when one prompt carries several subjects, a camera move, and a sound cue at once, and leave it off for a single action.

  1. 01

    Subject

    Name who or what is in frame and the one action they take. Two competing actions is where Wan 3.0 briefs start to drift.

  2. 02

    Camera

    One move: push in, orbit, handheld follow, or locked off. Add the lens character you want in the same clause.

  3. 03

    Sound

    Say what should be audible and what causes it. Unnamed audio still arrives, but it is whatever the scene suggests.

  4. 04

    Delivery

    Seconds, aspect ratio, and resolution. Duration and resolution are the only two fields that move the price.

Example Wan 3.0 brief16:9 · 5 seconds · audio on

The potter keeps turning the half-thrown bowl, clay dust drifting through the window light, apron shifting slightly. Slow push-in, quiet workshop room tone, no music.

Before committing a long render, compare this exact model against Gemini Omni 1.1 Flashfor mixed-reference editing and a 4K ceiling, or review H3 Max vs Wan 3.0when rapid prompt iteration is the deciding constraint. A vertical episode that needs a character bible and several Wan 3.0 shots belongs on the AI short drama generator. Directed camera cuts inside one 15-second file belong on the multi-shot AI video generator. Ranking Wan 3.0 against other catalog rows for a paid spot is the best AI video model for ads page.

Wan 3.0 FAQ

What is Wan 3.0?

Wan 3.0 is Alibaba's newest Wan video model, released through inference APIs on 24 August 2026. It generates 2 to 30 seconds of video with synchronized audio from a text prompt, a still image, or reference media, at 480p, 720p, or 1080p.

What is the difference between Wan 3.0 and Wan 3.0 Prime?

Both tiers expose the same three modes, the same 2–30 second range, the same resolutions, and the same audio. Prime is the premium tier and costs about 40% more per second — $0.14 against $0.10 at 720p. Run the cheaper tier while you are still deciding the shot, then re-run the winner on Prime.

How much does Wan 3.0 cost on Voor AI?

Wan 3.0 bills per second of finished video: $0.05 at 480p, $0.10 at 720p, and $0.20 at 1080p. Wan 3.0 Prime bills $0.068, $0.14, and $0.28 for the same resolutions. A five-second 720p clip is therefore $0.50 on Wan 3.0 and $0.70 on Prime, and the generator shows the credit estimate before you spend anything.

Which Wan 3.0 modes are available here?

All six: text-to-video, image-to-video, and reference-to-video on both the Wan 3.0 and the Wan 3.0 Prime tier. This page opens on Wan 3.0 text-to-video; switch modes in the model picker without leaving the route.

Does Wan 3.0 generate audio?

Yes. Audio is on by default and comes out of the same render as the picture, so a scored or ambient bed arrives synchronized rather than added afterwards. Turn the audio switch off if you plan to lay your own track over the clip.

How long can a Wan 3.0 clip be?

Two to thirty seconds, set as a whole number. Duration is the price lever: cost scales linearly with seconds, so a 30-second 1080p Wan 3.0 render costs fifteen times a 2-second one.

What does thinking mode do in Wan 3.0?

Deep-thinking mode lets the model reason over a complicated prompt before it starts rendering, which helps when a brief names several subjects, a camera move, and a sound cue at once. It is off by default and adds latency, so leave it off for simple single-action shots.

How does Wan 3.0 reference-to-video handle a reference image?

It rebuilds a scene in the reference's likeness rather than continuing that exact frame. In our demo it kept the potter, the apron, the bench, the bowl, and the dusty palette but replaced the window wall with a shelf of finished pots. Use image-to-video when the first frame must be preserved exactly, and reference-to-video when you want the look and the identity carried into a new setup.

Can I reuse a seed between Wan 3.0 and Wan 3.0 Prime?

No — a seed is only comparable inside one tier. We ran the same prompt at seed 424242 on both and got two different compositions of the same scene, so treat a tier switch as a fresh generation rather than a quality upgrade of the clip you already have.

Generate with Wan 3.0

Write a brief, pick a length and a resolution, and see the credit estimate before Wan 3.0 renders anything.

Start creating