1. Upload a still that survives motion
Faces need sharp eyes and even light for podcast image audio to video. Product packs need readable labels — blur amplifies when motion starts. Landscapes need horizon detail so cloud drift does not smear into mush. Square crops work; 9:16 stills map cleaner to Reels.
2. Describe one restrained motion beat
Ask for one subject action and one camera move. The audio track does not time those actions, so describe the visible beat in the prompt instead of relying on the music or voiceover.
3. Generate with optional background audio
Upload the still, prompt, and optional track, then review with sound on. If you need a dedicated dialogue mouth tool, browse Lip Sync.
Limits and failure boundaries we hit in tests
Choose 5, 10, or 15 seconds. Audio longer than the selected video is truncated; shorter audio leaves silence at the end. If the track must drive a speaking subject, use Create Video from Audio AI instead.
Export checklist after generation
Confirm aspect matches the destination (9:16 for Reels/TikTok, 16:9 for YouTube intros). Watch once with sound on for sync and loudness. Import the original still as a color reference if you grade further. If the bed is longer than one clip, trim the audio or generate a second pass with a related still — viewers notice identical loops under voiceover.
When to use an AI route instead
Use this page when the image and written motion brief are primary and audio is only a bed. Use Create Video from Audio AI for speech-driven faces, or Photo to Video with Music for a local multi-photo slideshow.