Text and image input
Start with a text prompt, upload a reference image, or combine both. The model interprets brief and reference together rather than treating them as separate instructions.
Motion reference · two different stresses
In the horse clip, watch whether identity holds through the movement. In the ski clip, check whether camera, snow, scale, and the ending all hold together.
Start with a text prompt, upload a reference image, or combine both. The model interprets brief and reference together rather than treating them as separate instructions.
Dialogue, ambience, and effects are generated alongside the video, not layered on after. Connect each sound to a visible cause in the scene.
FLUX 3 Video supports clips up to twenty seconds. That length needs an arc — an opening state, one cause-and-effect event, and a defined ending.
A longer clip earns its length through progression, not more adjectives.
Subject, world, camera, room tone
One action starts and the world responds
Camera exposes scale or consequence
Action completes; final frame holds
0:00 — matte-black box on a concrete plinth, one overhead softbox. 0:04 — lid lifts; mechanical click and a short pneumatic hiss. 0:12 — camera orbits right at plinth height while the product catches a rim light. 0:20 — settle front-on with negative space above; hold one second.
Length and ratio decide what the shot can contain. Changing them afterwards means rewriting.
| Setting | Options | What it changes |
|---|---|---|
| Duration | 5–20s, one-second steps | Cost scales with length. Past about twelve seconds the clip needs a real arc, not a longer hold. |
| Resolution | 720p or 1080p | 1080p costs noticeably more per second. Draft at 720p, promote the take that works. |
| Aspect ratio | 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, 9:16 | Decides the motion corridor and where overlaid text can sit without covering the subject. |
| Audio | On by default | Generated in the same pass as the picture, so effects land in sync with what causes them. |
| Mode | Full or draft | Draft runs the same brief for roughly a third of the credits. Use it to find the take. |
| Input | Text, or image plus text | Image-to-video holds the source composition and adds motion from your description. |
The credit estimate in the generator updates as you change these, so you can price a twenty-second 1080p shot against three draft passes before committing to either. For a first attempt at an unfamiliar subject, drafts usually win: you learn whether the action reads at all for a fraction of what one wrong full render costs.
Motion benchmark board
Run one test at a time. A controlled benchmark tells you whether the model handled material, contact, scale, speech, or space—not merely whether the clip looked exciting.
A dancer turns against a painted wall; the skirt lags behind the hips, catches, and settles. Watch whether fabric keeps a consistent weight through the turn instead of snapping to the new pose.
Fireworks over water above a silhouetted crowd. The test is whether sparks fall on their own arcs, the water carries the burst, and faces stay dark without dissolving into noise.
Weathered hands pull a rope knot tight on a boat gunwale. Hands are where video models fail most visibly — count fingers, follow the rope through the knot, and check the grip never swaps.
A slow drift along a ramen counter on a rainy night. Stools, bowls, and the cook hold their positions relative to each other as the camera travels, and reflections move with it.
Voor workflow tutorial · not a FLUX 3 demo
The interface walkthrough shows where a text-to-video brief, model choice, settings, and generated result live in Voor. Use the active model label as the truth for today’s render, then keep the same prompt, reference roles, ratio, duration, and review notes ready for a future FLUX 3 Video endpoint.
Read the action aloud against the timeline. Remove any event that cannot physically complete before the resolution window.
Assign identity, material, environment, or motion to each file. Unlabelled evidence creates competition instead of control.
Name a cause for every important sound and protect dialogue with a restrained ambience bed. Silence can be an intentional beat.
Inspect the whole clip for geometry, contact, direction, reflections, and a settled final frame. Novelty never replaces editorial usability.
FLUX 3 Video FAQ
Duration, audio, resolution, draft mode, and what belongs in a brief that survives twenty seconds of motion.
FLUX 3 Video is Black Forest Labs' video model. It generates clips of 5 to 20 seconds with native audio, working from a text prompt, a reference image, or both. On Voor it runs through the generator on this page — no API key and no separate account required. The model handles text-to-video and image-to-video in the same interface, and the audio track is generated alongside the picture rather than added in a separate pass. That single-pass design means environmental sounds, dialogue, and effects land in sync with the visible action by default, which removes the manual alignment step that makes layered audio workflows tedious.
Yes, and audio is on by default. Dialogue, ambience, and effects are produced in the same pass as the image rather than layered on after, so they land in sync with visible action. Write sound into the brief the way you write physical action: name the source, quote any spoken line, and tie each effect to something happening on screen. The model does not invent a musical score unless you ask for one, so if you want a clean sound bed of just ambience and foley, simply leave music out of the prompt. Silence is also a valid choice — stating 'no ambient sound in the final second' gives you a clean audio tail for editing.
Anywhere from 5 to 20 seconds, in one-second steps. Cost scales with duration, so pick the length from the edit you actually need rather than defaulting to the maximum. Twenty seconds only pays off when the clip has somewhere to go — an opening state, one cause-and-effect event, and a defined ending that an editor can cut on. Past about twelve seconds, the brief needs a real arc with progression; a static scene that simply holds for twenty seconds wastes the extra duration and the extra credits. Draft at a shorter duration first to confirm the action reads, then extend once you know the composition works.
720p or 1080p, in 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, or 9:16 — or auto, which lets the model choose. Set the ratio before writing the brief: it decides where motion can happen and where overlaid text will sit, and changing it after a render means starting over because composition is not ratio-independent. 1080p costs noticeably more per second than 720p, so the practical workflow is to iterate at 720p until the action, framing, and sound all work, then promote the winning take to 1080p for the final render. That approach typically costs less than getting a single 1080p generation wrong twice.
Draft renders the same brief at 720p for roughly a third of the credit cost. Use it to test whether an idea holds up — framing, pacing, whether the action reads, whether the sound cues land where they should — then rerun the winning brief at full quality. Iterating on drafts and promoting once is almost always cheaper than getting a full render wrong twice. Draft mode is also useful for exploring unfamiliar subjects: if you have never prompted a specific type of motion before — glass shattering, fabric tearing, liquid pouring — a draft tells you whether the model can handle the physics at all before you commit full credits to the attempt.
Yes. Image-to-video takes a source image plus a prompt describing what should move. It works best when the subject is clear and the background is not competing for the model's attention. Describe the motion you want rather than re-describing the picture — the model can already see the image and will treat your text as instructions about movement, camera, and sound, not a replacement caption. Keep the source image clean and well-composed; artefacts, heavy compression, or ambiguous edges in the input tend to amplify during animation. Crop and retouch the still before uploading rather than hoping the motion pass will smooth things over.
The opening frame, one physical action with a clear cause and visible effect, camera position and any movement, the sound environment and its sources, and what the last frame looks like. Concrete physical verbs beat mood words: say 'rain beads on the jacket and the camera tracks at shoulder height' rather than 'it should feel cinematic'. End with a settling instruction — 'camera holds, subject faces left, one second of room tone' — so the clip has a clean tail for editing. Prompts that read like a shot list produce better results than prompts that read like a poem, because the model can execute specific spatial instructions and has to guess at abstract emotional ones.
Change one variable at a time. Keep the subject, duration, and closing frame fixed while you adjust camera distance, lighting direction, or a single audio element. Rewriting the whole brief between runs makes it impossible to tell what actually helped — and every run costs credits. Record the exact prompt and a specific reject reason for each failed take. The reject log is the most valuable tool in the iteration process: after three or four runs, patterns emerge — the model consistently drops an object, or lighting from a certain direction causes shadow artefacts — and those patterns tell you what to change next. Without the log, you are guessing on each attempt.
Yes — short product reveals, material studies, unboxing sequences, and social hooks all fit the model's strengths. Put product geometry, label placement, brand colour, and logo treatment in the continuity rules so they survive the motion. The generated footage still needs the same review as filmed footage: product accuracy, advertising claims, trademark clearance, likeness rights for any recognisable person, and whatever the destination channel requires for compliance. Keep the prompt, source references, generated output, and approval decision archived together with the published asset so you can trace provenance, regenerate a variant, or answer a rights inquiry months later.
Black Forest Labs' video model, generating up to 20 seconds with native audio from a text prompt or a reference image. Draft mode runs the same brief for about a third of the credits.