Image to Video
image_to_videoInput

Prompt
The dog turns its head and wags its tail in warm sunlight.
Output video
Image-to-video output — dog wagging tail with synced audio

The Omni Flash Generator on Voor AI runs Google's Gemini Omni Flash — a multimodal video model, not camera flash. Pick a mode in the bar above, upload a still or clip, write a prompt, and download MP4. Three FAL endpoints sit behind one interface: image-to-video, reference-to-video, and edit.
I2V
Image to Video
One still becomes a 3–10s clip with synced audio.
R2V
Reference to Video
Up to 10 reference images steer subject and style.
Edit
Edit
One instruction rewrites an existing MP4.
Google's Gemini API names three tasks — image_to_video, reference_to_video, and edit. The mode bar maps directly to them. Match your upload type to the task, paste a specific prompt, then reroll until motion and audio land.
STEP 1
Image to Video for one still. Reference to Video for storyboard or product refs. Edit for conversational changes on existing footage.
STEP 2
Name camera verbs and audio mood. Tag refs with <IMAGE_REF_0>. For edits, add “Keep everything else the same.”
STEP 3
Generate 3–10s takes with audio on I2V and R2V, download MP4, chain in your editor, or run Edit on the winner.
Each block is a verified run: the inputs and prompts below were sent to FAL's google/gemini-omni-flash endpoints; the output MP4 is the model's response — not a stock clip from the playground sidebar. Reproduce the same brief in the generator above.
image_to_videoInput

Prompt
The dog turns its head and wags its tail in warm sunlight.
Output video
Image-to-video output — dog wagging tail with synced audio
reference_to_videoInput


Prompt
A chimpanzee wearing overalls frolics in the grassy field, gently playing with the butterflies. In the background, a circus tent and carousel beckon.
Output video
Reference-to-video output — generated from the two reference stills above
editInput
Prompt
Make this video anime. Keep everything else the same.
Output video
Edit output — same scene after “Make this video anime” pass
Keyframes and mood plates you can upload into Image to Video, Reference to Video, or Edit. Copy the prompt, match the mode, generate.

I2V keyframe
The dog turns its head and wags its tail in warm sunlight — Omni Flash Generator image_to_video.

Edit source still
Make this video anime. Keep everything else the same — Omni Flash Generator edit.

Product I2V
Slow dolly-in on a matte wireless earbud case, soft studio rim light — Omni Flash Generator.

Scene plate
Cozy cafe at golden hour, slow dolly, inviting wide shot — Omni Flash Generator reference_to_video.

Poster mood
A futuristic city with neon lights and flying cars, cyberpunk style, 9:16 — Omni Flash Generator.

Subject lock
Gentle orbit around hero subject, warm studio light, synced ambient audio — Omni Flash Generator.
Google built Gemini Omni Flash as a natively multimodal stack — text, images, audio, and video in, high-resolution video with audio out. The Omni Flash Generator exposes that through three tasks, each with a distinct upload and prompt pattern.
Your still is the first frame. Google recommends specific motion language — “turns its head,” “slow dolly-in” — not vague “make it move.” Output includes synced audio when the scene implies it.
Reference images guide subject and style without becoming literal frames. Bind roles with <IMAGE_REF_n> tags in the prompt. Ideal for character plus product, storyboard frames, or mood boards.
Upload one source MP4 and describe a single change — anime restyle, remove an object, shift lighting. Add “Keep everything else the same.” for local edits. Voice editing is not supported.
01Flash is Google's model name — not camera flash. Write motion verbs, not lighting gear.
02Image to Video: high-res still + specific camera move — avoid vague “make it move” lines.
03Reference to Video: tag uploads with <IMAGE_REF_0> when you need role control.
04Edit: one short instruction per pass; add “Keep everything else the same.” for local changes.
Copy a line, pick the matching mode, upload, and generate. Reroll three times before you rewrite the brief — Google's field-tested patterns outperform generic motion requests across every Gemini Omni Flash task.
High-resolution still plus one camera verb. Your upload becomes the literal first frame.
Reference images guide subject and style — they are not literal first frames. The demo above uses FAL's bundled circus + meadow stills with the matching Veo 3.1 prompt. For your own refs, tag roles with <IMAGE_REF_n>.
One source MP4 per pass. Google's examples include invisible-object edits and mirror-ripple transforms.
The strongest Omni pages use abundant video and clear mode names. Voor adds the missing decision layer so source media, reference media, and edit media are not confused.
Use one literal first frame when composition and subject placement should carry into the generated shot.
Use several images as identity, product, or style guides and assign their roles in the prompt.
Upload an existing video and request one bounded transformation while preserving everything else.
Generation modes can return synced audio; visual edit passes should still be reviewed and remixed deliberately.
The Omni Flash Generator is a three-mode video workspace for Google Gemini Omni Flash. Pick Image to Video, Reference to Video, or Edit, upload stills or a clip, write a prompt, and download MP4 — no separate Google API key on Voor AI. Each mode maps to a dedicated google/gemini-omni-flash FAL endpoint with pricing shown before you generate.
No. Flash is the Google DeepMind model name for a multimodal video system — not photography flash, not a speed-tier label for still images. Google published the model card in May 2026; the name refers to a high-performance video stack with native audio output.
Google DeepMind describes it as a step toward models that create and edit anything from any input — starting with video. Capabilities on the roadmap include T2VA, I2VA, reference-to-video, video editing, and image generation. Today the Omni Flash Generator focuses on the three video tasks Google documents in the Gemini API.
Image to Video for one hero still — product packshots, portraits, or illustrations. Reference to Video when you need multiple anchors: cat plus yarn, character plus wardrobe, storyboard frames. Edit when you have footage and want a style or object change without reshooting, using short instructions Google validates in its editing examples.
Yes on Image to Video and Reference to Video. Google’s model card specifies high-resolution video with audio; FAL returns synced speech, ambience, or music when your prompt implies them. Edit changes visuals only — remix audio in post if the take needs a new mix.
3–10 seconds per pass on FAL, with 16:9 or 9:16 aspect ratio. Google Cloud preview documentation lists a 10-second maximum at 720p. Chain multiple passes in your editor for longer stories; treat each Omni Flash Generator render as one shot in a sequence.
Google’s Gemini API docs include: a marble on a chain-reaction track (text-to-video); drawing-to-realistic-footage (image-to-video); cat batting yarn with two reference images; and edits like “Make the violin invisible” or mirror-ripple transforms. Copy those patterns — specific motion beats outperform vague “make it move” lines.
Google’s model card states that maintaining complete consistency through edits, complex motion, and perfectly accurate text remains a challenge. Proofread every frame before paid placement; keep headlines short and plan manual fixes for logos or legal lines.
Upload reference images in order. In the prompt, bind them with <IMAGE_REF_0>, <IMAGE_REF_1>, and so on to assign character, product, or background roles — the same pattern FAL documents for google/gemini-omni-flash/reference-to-video. Google distinguishes source media (literal first frame) from reference media (style and subject guides).
Google’s Interactions API chains edits with previous_interaction_id so context carries across turns. On Voor AI, Edit mode takes one source MP4 plus a plain-language instruction per pass — e.g. “Make this anime. Keep everything else the same.” Reroll or chain clips manually until the edit sticks. Voice editing is not supported.
New accounts receive 3 welcome credits. Paid cost varies by duration and mode; check the estimate before batching variants.
The Omni Flash Generator on Voor AI runs google/gemini-omni-flash image-to-video, reference-to-video, and edit — Google's multimodal video model with synced audio on generation modes.
Pick Image to Video, Reference to Video, or Edit, upload your still or clip, paste an official-style prompt, reroll until the take works.
Generate with Omni Flash GeneratorGoogle Gemini Omni Flash on Voor AI — Image to Video, Reference to Video, and Edit in one bar. Upload, prompt with official patterns, download MP4 with audio.