Gemini Omni 1.1 Flash builds the picture and the sound in one pass
Most video models take a prompt and hand back silent footage. Gemini Omni 1.1 Flash reads the written direction, the frames you supply, short reference clips, the camera notes and the sound as one brief, then renders them together. The clip on this page came out of the same Text to Video mode loaded in the workspace above: one subject held through several poses, the studio lighting consistent the whole way through, and real production sound in the finished MP4.
Text to video10 seconds · 720p · native audio
3–10sWhole seconds, across the three generation modes
4KTop of four tiers, from early 360p checks to final 4K output
2 ratiosLandscape 16:9 or vertical 9:16, nothing in between
AudioDialogue, ambience, effects and music, rendered with the picture
Four ways into the same timeline
Let your source material pick the mode
The four Gemini Omni 1.1 Flash modes are not quality tiers. They answer four different questions: do you need to invent a shot, move a still, combine references, or fix footage you already have? Feed in the least material that still pins down what must not change.
Text to video
Fashion editorial in a warm-grey studio. Five deliberate poses, slow dolly, 85mm lens, softbox key, restrained film grain, quiet room tone and fabric movement.
Start here when all you have is the brief. Casting, set, camera, motion and sound all get invented together.
Image to video
Close vertical café selfie. He raises the cup, smiles naturally, and says one short line; keep age detail, window light, and the handheld phone texture.
Your still becomes frame one. Add an end frame only when you care more about where the shot lands than what happens on the way there.
Reference to video
Build a tense period-drama confrontation using the supplied character and location references, preserving wardrobe, faces, atmosphere, and the visual language of the set.
Several sources, each with a job to do: who is in the shot, where it happens, how the movement should feel.
Edit video
Replace the bottle with an apple. Preserve the person, hand motion, camera, lighting, shelves, timing, and every part of the clip not named in the change.
Upload a finished clip and change one thing. Saying what to leave alone is what keeps the rest of the frame from drifting.
Mode decision board
What you upload changes how Gemini Omni 1.1 Flash behaves
No media · Text to Video
Go prompt-only when casting, art direction and framing are all still open. The model will stage a believable scene out of general world knowledge, but nothing holds it to a person or a product you never showed it. Describe what the camera would actually see instead of writing a mood slogan. And if one specific face or product has to be right, switch to a reference workflow now rather than after three text-only rounds.
One or two frames · Image to Video
The first image is literally frame one, which makes this the mode for an approved packshot, portrait, illustration or storyboard panel. An optional end image pins down where the shot finishes. Just keep the two frames compatible: same subject, camera geometry that could plausibly connect them, and a change that could really happen in ten seconds. For other still-first options, compare the models in Image to Video.
Mixed evidence · Reference to Video
Reference mode takes up to ten images and up to three reference videos, none longer than three seconds. Ten loosely related mood shots do worse than four that each do a specific job, so tell the prompt what every item is for: character, wardrobe, product, location, motion or voice. The model reads across all of them, and where a role is vague it will happily blend two things you meant to keep apart.
Existing timeline · Edit Video
Edit mode inherits the source clip's timing and composition. Ask for one change — swap an object, restyle the world, change the weather, replace a background — and say out loud what has to stay. There are no masks, keyframes or undo history here, so download every pass you are happy with before asking for the next one. If the job is really a full visual restyle,AI Reimagine for Videos is the better fit.
Conversation as an edit list
Treat every Gemini Omni 1.1 Flash revision as a single edit note
Editing by chat feels casual, and that is the trap. Every turn should produce one file you can hold up against the last one. “Improve this video” leaves you no way to tell what actually changed; “replace the bottle with an apple” gives you something to check. The pass below keeps that trail clean while you work through a clip.
SourceLock the accepted camera and timing
Upload the take you are happiest with, not a rough one you already plan to redo.
Turn 01Name one visible change
“Replace the bottle with an apple; preserve the hand movement and lighting.”
ReviewWatch picture and sound end to end
Check the edit boundary first, then identity, background, dialogue, ambience and the last frame.
Turn 02Continue only from an accepted result
Save the file you approved, then ask for the next single change in a fresh pass.
A prompt that can be reviewed
Five layers for a Gemini Omni 1.1 Flash brief
Keep subject, action, camera, sound and output settings in separate sentences and you can rewrite one of them without disturbing the other four. The same five layers work for generating and for editing; the only difference is whether the first one describes something you imagined or something you uploaded.
01
Continuity anchor
Name the person, object, wardrobe, place or composition that is not allowed to drift. In edit mode, name the change and the things that have to survive it.
02
Visible action
Give each subject one readable action with a beginning and an end. A short clip can hold a turn, a reveal, an exchange or a camera move. It cannot carry a whole campaign script.
03
Camera sentence
Pick one main move and one lens character. “Slow push-in on an 85mm portrait” is a direction. “Cinematic camera” is just an adjective.
04
Sound sentence
Write the dialogue word for word, then handle room tone, effects and music in their own clauses. Audio behaves best when every sound has something on screen, or in the story, causing it.
05
Output settings
Set duration, orientation and resolution in the controls, not buried in the prose. The prompt is for intent, the fields are for numbers.
A full image-to-video briefPreserve the person, glasses, charcoal sweater, café layout, and soft window light from the first frame. He raises the cup once, pauses, smiles with a small eye movement, and says: “Third one today. Don't tell my doctor.” Handheld phone framing with a subtle natural drift, no portrait-mode edge blur. Quiet café room tone, cup contact on the table, speech clear above the background. Eight seconds, vertical 9:16.
Delivery gate
Approve the whole clip, not a favorite frame
Gemini Omni 1.1 Flash can nail the poster frame and still come apart four seconds in: a hand, a reflection, a spoken line, something moving in the background. Watch each candidate at normal speed, once muted and once with sound, then scrub the first and last second, which is where continuity usually breaks. A take is done when the picture, the audio and the handoff to the next shot all hold up.
Identity: the face, product shape, wardrobe and distinctive marks survive the motion.
Camera: the move you asked for stays readable, with no surprise cut or extra orbit.
Sound: dialogue is clear, effects have visible causes, and the ambience does not jump at the end.
Frame edges: in both crops, the action you care about stays clear of caption and UI zones.
Sequence: the first and last frames can actually cut against the shots on either side.
Need something else? The longer-duration and specialist models sit inVideo Models, and if the still is not right yet, build it in Nano Banana 2 first, then come back to Gemini Omni 1.1 Flash for the motion.
Gemini Omni 1.1 Flash FAQ
What is Gemini Omni 1.1 Flash?
Gemini Omni 1.1 Flash is Google's stable multimodal video model released in August 2026. It accepts text, images, and short videos, then returns a 3–10 second video. It works four ways: text-to-video, image-to-video, reference-to-video, and editing an existing clip in plain language.
Can Gemini Omni 1.1 Flash generate audio?
Yes. The three generation modes create synchronized native audio with the video. Describe dialogue, ambience, music, and sound effects in the same prompt as the picture. Always review with sound on, because a take that looks great can still need a different audio direction.
Which Gemini Omni 1.1 Flash mode should I choose?
Choose Text to Video when you only have a written brief, Image to Video when an approved still must become frame one, Reference to Video when several images or short clips define identity and setting, and Edit Video when you want to change an existing clip without rebuilding everything.
How long can a Gemini Omni 1.1 Flash video be?
The generation modes accept any whole-second duration from 3 through 10 seconds. The edit mode follows the length of the uploaded source clip. For anything longer, plan it as separate shots with clear continuity notes and assemble the approved clips in an editor.
What resolutions does Gemini Omni 1.1 Flash support?
There are four tiers: 360p, 720p, 1080p, and 4K. Use 360p for early direction checks, 720p to judge motion and audio, 1080p for most final social work, and 4K once the shot direction is locked.
Can Gemini Omni 1.1 Flash use an end frame?
Yes. Image to Video accepts a required first image and an optional end image. The pair constrains where the shot starts and finishes, but the model still invents the motion between them, so choose frames with compatible subjects, camera position, and lighting.
How do reference images and reference videos work?
Reference to Video accepts up to ten images and up to three reference videos. Each reference video must be no longer than three seconds. Say in the prompt what each reference is for, so the model knows which one defines the character, the object, the setting, the movement, or the voice.
How is Gemini Omni 1.1 Flash different from the preview model?
Gemini Omni 1.1 Flash is the stable successor to the earlier Gemini Omni Flash preview. The stable release adds a clear text-to-video workflow alongside image, reference, and edit modes, exposes 360p through 4K output, and is the version to choose for new work.
Create with Gemini Omni 1.1 Flash
Pick the mode that fits what you are starting from, set duration and resolution, then review your settings before the render starts.