Pick a ready voice, clone your own, or bring finished audio. Then turn one portrait into a short-form talking video.
Tutorial
Press play to hear the result
Portrait, audio, motion — all visible here.
Tap a voice to hear its ready-made sample instantly.
Portrait
Audio
Video
These are complete, playable examples rather than silent posters. Listen for the voice, then watch the mouth timing and small head movements at phone size.
The fastest workflow is not one giant Generate button. It is three small approvals in the right order, so the expensive render starts only after the source and voice are ready.
01
Choose a portrait with a visible mouth and stable lighting. A shoulder-up crop gives the model enough context for small head and upper-body motion without making the face too small.
INPUT / JPG PNG WEBP
02
Write for listening: one idea per sentence, punctuation where you want pauses, and a direct first line. Generate audio, then correct names, pacing, or energy before you animate anything.
CHECK / NAMES PACE TONE
03
Send the approved audio and portrait to Omni Human. Keep the direction simple and consistent with the source photo, then download the result for captions, B-roll, music, and brand finishing in your editor.
OUTPUT / MP4 WITH AUDIO
0–2 sec
Start with the problem, surprise, or outcome. Skip greetings and channel introductions unless the creator’s identity is the hook.
2–10 sec
Name the action, product behavior, or result the viewer can understand immediately. One specific benefit sounds more credible than a stack of adjectives.
Final line
End with a short CTA or conclusion, then add punctuation. A deliberate final pause is easier to cut than speech that runs into the last frame.
Use your own portrait or obtain specific permission from the person shown. The permission should cover synthetic animation, the script, where the clip will appear, whether it is advertising, and how long the content may be used. A photo being public does not grant those rights.
Do not present an avatar as a genuine customer testimonial, celebrity endorsement, news report, or personal message when it is not. Keep the approved script, source image, voice choice, consent record, and final export together. Follow the disclosure rules of the platform, client, market, and ad network where you publish.
Upload a clear portrait, write the exact line you want spoken, choose a voice, and generate a voice preview. After you approve the audio, Voor sends the portrait and approved track to the selected motion engine to create the talking-avatar video.
Pronunciation, pacing, and tone are much faster to judge in an audio preview. Approving the voice first helps you catch script problems before the more expensive motion render starts.
Use one clearly visible person, a front-facing or lightly angled head, an unobstructed mouth, even light, and enough room around the shoulders. Avoid tiny faces, heavy motion blur, hands covering the mouth, and large sunglasses.
The selected motion engine sets the limit. Omni Human supports short clips up to about 15 seconds, while Omni Human 1.5 accepts approved audio under 35 seconds. The studio shows the active limit before rendering.
Yes. Choose the language and a compatible voice, then listen to the complete preview. Names, prices, acronyms, and mixed-language phrases should be checked by a fluent speaker before publishing.
You can create hooks, product explainers, localized variants, and creator-style drafts when you have the necessary image, script, voice, brand, and advertising rights. Do not fabricate testimonials or imply that a real person endorsed a product without permission.
Only when that person has given informed permission for this use. A public photo is not automatic consent. Do not use the tool for deceptive impersonation, fraud, harassment, or undisclosed endorsements.
Paid customers are private by default. Generations made with welcome credits may be eligible to appear in Explore. The studio states the active visibility behavior before you generate.
Generate the clean avatar take here. Use the adjacent tools when you already have audio, need a standalone voice track, or want to add motion that is not speech-led.