“Cheerful and upbeat, like a friendly station announcer.”
4.9 s · 2 credits
Gemini TTS reads your script in 30 prebuilt voices, follows a plain-language style instruction, and can voice a two-speaker dialogue in one file.

Real Gemini 3.8 Flash TTS output · 24.5 seconds · 6 credits
One request, two people. This podcast cold open was written as six lines of Name: text, and Gemini TTS returned a single file that switches between Maya and Theo on every turn, with a laugh where the script asked for one.
MayaOkay, we're recording. Welcome back to Small Kitchen, Big Ideas.
TheoToday's topic is the humble pancake, and I have strong opinions.
MayaYou always have strong opinions. <laugh> Go on.
TheoRest the batter for ten minutes. That's it. That's the whole secret.
MayaTen minutes? That's the big reveal?
TheoTrust me. Lighter, fluffier, every single time.
Style instruction for the whole take: “Relaxed, warm podcast banter between two old friends.”
Same voice, same words, three directions
Every take below reads the same line in the Sulafat voice. Only the style instruction changed. Pacing moved with it: the tired reading is two seconds longer than the cheerful one.
“The last train leaves at midnight. If you miss it, you're walking home.”
“Cheerful and upbeat, like a friendly station announcer.”
4.9 s · 2 credits
“Whispered and conspiratorial, as if sharing a secret.”
5.5 s · 2 credits
“Exhausted and deadpan, at the end of a very long night shift.”
6.9 s · 2 credits
30 prebuilt voices
Google describes each prebuilt voice in one word. Use that word to shortlist two or three, then let the style instruction do the rest. Highlighted voices are the ones you can hear on this page.
No language setting needed
Gemini TTS reads the language it is given. Both announcements use the same prebuilt voices as the English takes; the translation under each script is for you, not for the model.
Bienvenue au musée. Dans cette salle, vous découvrirez des cartes marines dessinées à la main il y a plus de trois siècles. Prenez votre temps : chaque détail raconte un voyage.
Welcome to the museum. In this room you will find sea charts drawn by hand more than three centuries ago. Take your time: every detail tells of a voyage.
Style: Calm, knowledgeable museum guide, unhurried.
本日はご来店いただき、ありがとうございます。ただいま、焼きたてのクロワッサンをご用意しております。どうぞごゆっくりお選びください。
Thank you for visiting us today. Freshly baked croissants are ready now. Please take your time choosing.
Style: Bright, polite in-store announcement.
Flash or Flash Lite
We sent the same lesson, voice Kore and instruction “Clear, encouraging cooking-class instructor” to both tiers. Both read every step in order, and the lengths differ by one second. Flash Lite costs about a third less, which makes it the sensible default for drafts and high-volume notification audio.
Step one: heat the oven to two hundred degrees. Step two: while it warms, cut the vegetables into even pieces, so everything roasts at the same speed. Step three: a little oil, a pinch of salt, and into the oven they go.
| Measure | Flash | Flash Lite |
|---|---|---|
| This 220-character lesson | 4 credits | 3 credits |
| Length of the take | 16.4 s | 17.4 s |
| 100 characters | 2 credits | 2 credits |
| 1,000 characters | ≈ 16 credits | ≈ 11 credits |
| 5,000 characters (maximum) | ≈ 79 credits | ≈ 53 credits |
Credits count the spoken text only; style instructions are free. New accounts get 50 free welcome credits, about two dozen short Gemini TTS lines.
The form behind the dialogue
These are the five fields in the generator above, filled in exactly as they were for the podcast sample.
Maya: Okay, we're recording…The exact words. Up to 5,000 characters; <laugh> and <sigh> work inline.Relaxed, warm podcast banter…How to speak, never spoken. One sentence about mood, pace and listener.KoreThe single voice, or the first name that appears in a dialogue.onReads Name: text lines as a conversation. Exactly two names.PuckUsed for the second name. Ignored when dialogue is off.Gemini TTS is Google's text to speech built on its Gemini models. On Voor it runs as two tiers, Gemini 3.8 Flash TTS and Gemini 3.8 Flash Lite TTS. Both read your exact words in one of 30 prebuilt voices, follow a plain-language style instruction, and can voice a two-person dialogue in a single file.
Yes. Turn on Two-speaker dialogue, write every line as Name: text using exactly two names, and pick a voice for each person. Gemini TTS keeps the turn order and switches voices for you. Three or more speakers are not supported in one take.
The style instruction describes how to speak, for example "whispered and conspiratorial" or "calm museum guide, unhurried". It is not read aloud and it applies to the whole take. Keep it to one short sentence about mood, pace and audience.
Both tiers share the same 30 voices, style instructions and dialogue support. Gemini 3.8 Flash Lite TTS costs about a third fewer credits. On our cooking-lesson script the two takes were close in pace, 16.4 seconds on Flash and 17.4 on Flash Lite.
Credits follow the length of the spoken text. Flash is about 16 credits per 1,000 characters and Flash Lite about 11. A 100-character line costs 2 credits on either tier, and a full 5,000-character script about 79 on Flash or 53 on Flash Lite. Style instructions are not counted.
Gemini TTS detects the language from the script, so no language setting is needed. This page includes English, French and Japanese takes generated with the same prebuilt voices. Have a fluent listener approve any language you cannot judge yourself.
Yes. Write <laugh> or <sigh> inline where the sound should happen, as in the podcast sample on this page. Use them sparingly; a vocal event in every line stops sounding natural.
No. Gemini TTS uses prebuilt voices such as Kore, Puck, Charon and Sulafat. If you need speech in your own voice from a reference recording, use a model with voice cloning on the text to speech page, and only with that speaker's consent.
Voor accepts up to 5,000 characters per Gemini TTS request, around five minutes of speech. Split longer narration into scenes, keep the same voice and style instruction, and join the files in an editor.
Start with a short exchange, pick a voice for each name, and add one sentence of direction. Keep the take that sounds like a real conversation.
Generate with Gemini TTSOptional cookies help us understand usage and measure ads. Essential cookies stay on so Voor works.