Preset voice
Choose one of nine exposed speakers—Aiden, Dylan, Eric, Ono_anna, Ryan, Serena, Sohee, Uncle_fu, or Vivian—then provide text, language, and an optional style instruction. This is the lowest-friction route for narration and prototypes.
Hello, I'm Aiden and it's very nice to meet you.
Both WAV files are preserved on the Voor CDN. One demonstrates a preset-voice baseline; the other demonstrates how a cloned-voice take should be reviewed for texture, stability, and permission.
custom_voice · English · Aiden
“Hello, I'm Aiden and it's very nice to meet you.”
Listen for the initial greeting, consonant clarity, pause after the name, and whether the final phrase lands naturally. Use the same checks on a short test before producing a longer script.
voice_clone · official repository sample
“Official Qwen3-TTS repository output demonstrating the voice-cloning path.”
Use this to judge voice texture and stability, not to infer permission. In production, the reference speaker must explicitly authorize the clone and its intended use.
Choose the mode before writing direction. Preset, clone, and design do not ask for the same inputs or carry the same permissions.
Choose one of nine exposed speakers—Aiden, Dylan, Eric, Ono_anna, Ryan, Serena, Sohee, Uncle_fu, or Vivian—then provide text, language, and an optional style instruction. This is the lowest-friction route for narration and prototypes.
Upload authorized reference audio and, preferably, its transcript. Qwen describes cloning from a short sample and cross-language use, but clean recording and explicit speaker consent remain part of the input contract.
Describe a new voice in natural language instead of selecting a real person: age range, vocal weight, pace, accent, energy, and context. Avoid asking for a named person’s imitation.
The connected schema lists Chinese, English, Japanese, Korean, French, German, Italian, Spanish, Portuguese, and Russian, plus automatic detection. A fluent listener still needs to approve pronunciation and locale.
The style instruction can direct pace and emotion. One coherent performance—calm tutorial, warm storyteller, urgent dispatch—works better than contradictory adjective stacks.
A name, a number, and the most emotional sentence expose pronunciation, pacing, and voice fit before a full chapter.
Fast narration with one of nine exposed speakers.
Authorized reference audio for a consistent voice texture.
Describe a new voice without imitating a named person.
Generate in the final language, then use a fluent reviewer.
Breaths, clicks, room noise
Names, numbers, consonants
Pauses, subtitles, music
Meaning, dialect, consent
The capability summary reflects the fields exposed by Voor’s live Qwen3 TTS generator: preset speakers, 10 listed languages, voice design, style control, and reference-led voice cloning.
Model capability is not delivery approval. Generated speech can mispronounce names, flatten emotion, or drift across long passages. Voice cloning also creates a consent and disclosure obligation that no quality benchmark can replace.
Yes. The generator is locked to the priced Qwen3 TTS model in Voor and belongs to the live text-to-speech collection.
Qwen3 TTS can synthesize speech from text, direct a voice style, and support reference-led voice or style workflows exposed by the connected model form.
Use the languages exposed by the current model form and write the script in its final language. Always ask a fluent speaker to review names, numbers, tone, and pronunciation.
The model supports reference-led voice workflows. Upload only your own voice or a recording whose speaker gave informed permission for this specific use.
Describe one coherent delivery such as warm documentary narration, restrained excitement, or calm customer support. Include pace, energy, pauses, and audience instead of stacking contradictory emotions.
Use short paragraphs, normal punctuation, phonetic help for unusual names, and separate takes for sharply different characters or moods. Write for listening rather than copying dense page prose.
The live generator calculates credits from the connected endpoint and the submitted text. Review the displayed estimate before generating, especially for long scripts.
Check your Voor plan, model terms, script rights, music or trademark references, and voice consent. Disclose synthetic voice where a platform, contract, or audience context requires it.
Proof-listen on headphones and a phone speaker for pronunciation, missing words, clipped consonants, breaths, room artifacts, emotional consistency, and appropriate loudness.
Choose a production preset, paste the final script, direct the delivery, and proof-listen before publishing.
Open Qwen3 TTS