Gemini 2.5 Flash TTS is Google’s text-to-speech model for turning writing into natural spoken audio, with 30 voices and support for 87 languages. It reads a script in one voice, or performs a conversation between two, and you direct the delivery by describing what you want in ordinary words.
That last part is the thing worth understanding. There are no sliders to balance and no markup language to learn. You write an instruction, the way you would brief a voice actor, and the model performs it.
Gemini 2.5 Flash TTS at a glance
| Model ID | gemini-2.5-flash-tts |
| Best for | Fast, low-cost everyday voiceovers |
| Input and output | Text in, audio out |
| Voices | 30 |
| Languages | 87, including 24 fully released |
| Speakers | One voice, or two for a conversation |
| Audio formats | MP3, WAV, OGG and PCM |
Tell it how to read, not just what to read
Two separate things go in. The script holds the words you want spoken. The instruction describes how to say them, and it never gets read aloud. That split is why the model takes direction so well, because your stage notes stay out of the script and the voice cannot read your instructions back to you.
Writing “say this in a curious way” ahead of your script is enough to change the whole read. Warm and unhurried, brisk and professional, gentle like a bedtime story: describe the performance and the model works toward it. This is also where accents come from, so asking for a regional accent in the instruction is what produces one. Changing the language setting on its own will not.
Directing pace, tone and emotion
For a change partway through a line, drop a cue in square brackets exactly where the shift should happen. Writing [extremely fast] in front of a legal disclaimer races through it the way a real voice actor would. [whispers] drops the delivery to a hush mid-sentence.
Three kinds of direction do most of the work:
- Tone and style. Ask for a mood, a register or a whisper, and the delivery follows.
- Pace. Speeding up or slowing down also sharpens pronunciation, which helps with tricky names and numbers.
- Accent. Name the accent you want in the instruction rather than relying on the language setting.
The model handles poetry, news copy and storytelling convincingly, and performs a specific emotion when you name one.
Language coverage
87 languages and regional variants are supported, and 24 of those are fully released: Arabic, Bengali, Dutch, English for both India and the United States, French, German, Hindi, Indonesian, Italian, Japanese, Korean, Marathi, Polish, Portuguese, Romanian, Russian, Spanish, Tamil, Telugu, Thai, Turkish, Ukrainian and Vietnamese.
The other 63 are still being refined, and that group covers British and Australian English, Canadian French, Mexican Spanish, European Portuguese and Mandarin Chinese, alongside a long tail running from Afrikaans to Urdu. They work today, though the output can shift as Google improves them, so build a campaign on the fully released list wherever you can.
Two-speaker dialogue
Label each speaker in the script, assign a voice to each name, and one generation produces the whole exchange. Writing “Sam:” and “Bob:” at the start of each line, then pointing Sam at Kore and Bob at Charon, gives you a two-person conversation in a single pass.
Generating each speaker separately and stitching the takes together afterwards loses the thing that makes a conversation sound real. Speakers in the same generation share context, so the timing, the reactions and the handoffs all land naturally. Keep it in one pass.
How long a script can be
The script box holds up to 5,000 characters, which is roughly five minutes of speech. That covers most voiceover work: a product explainer, a reel narration, a short episode. Longer projects are better split into separate generations anyway, since a single unbroken read gives you nothing to edit around if one line needs redoing.
When you do split a script, break at paragraph ends rather than mid-sentence, and keep every part on the same voice. A fragment that starts mid-clause has no context to set its tone, and a voice change between parts is audible immediately.
Generating a voiceover from here takes a script, an instruction and a voice.
Get answers to common questions
Gemini 2.5 Flash TTS is Google’s text-to-speech model for generating natural spoken audio from writing. It offers 30 voices across 87 languages, reads in one voice or two, and takes direction on tone, pace and accent in plain language.
Start generating voiceovers
Gemini 2.5 Flash TTS earns its place for the everyday work: narration for an explainer, a voiceover for a reel, a two-person script for a product walkthrough. Open the Picsart AI Playground, switch to Audio mode, write your direction ahead of your script, and pick a voice that suits the job.