Gemini 2.5 Flash TTS is Google’s text to speech model for turning writing into natural spoken audio, with 30 voices and support for 87 languages. It reads a script in one voice, or performs a conversation between two, and you direct the delivery by describing what you want in ordinary words.

That last part is the thing worth understanding. There are no sliders to balance and no markup language to learn. You write an instruction, the way you would brief a voice actor, and the model performs it.

Gemini 2.5 Flash TTS at a glance

Model ID gemini-2.5-flash-tts
Best for Fast, low-cost everyday voiceovers
Input and output Text in, audio out
Voices 30, being 14 female and 16 male
Languages 87, including 24 fully released
Speakers One voice, or two for a conversation
Audio formats MP3, WAV, OGG and PCM
Script length in Picsart Up to 5,000 characters

Tell it how to read, not just what to read

Two separate things go in. The script holds the words you want spoken. The instruction describes how to say them, and it never gets read aloud. That split is why the model takes direction so well, because your stage notes stay out of the script and the voice cannot read your instructions back to you.

Writing “say this in a curious way” ahead of your script is enough to change the whole read. Warm and unhurried, brisk and professional, gentle like a bedtime story: describe the performance and the model works toward it. This is also where accents come from, so asking for a regional accent in the instruction is what produces one. Changing the language setting on its own will not.

Directing pace, tone and emotion

For a change partway through a line, drop a cue in square brackets exactly where the shift should happen. Writing [extremely fast] in front of a legal disclaimer races through it the way a real voice actor would. [whispers] drops the delivery to a hush mid-sentence.

Three kinds of direction do most of the work:

  • Tone and style. Ask for a mood, a register or a whisper, and the delivery follows.
  • Pace. Speeding up or slowing down also sharpens pronunciation, which helps with tricky names and numbers.
  • Accent. Name the accent you want in the instruction rather than relying on the language setting.

The model handles poetry, news copy and storytelling convincingly, and performs a specific emotion when you name one.

Picking a voice

Thirty voices is a lot to work through one by one, so start from the job rather than the name. The names come from stars and mythology and tell you nothing about how any of them sound.

Everyday narration

Kore, Zephyr, Achird, Sadaltager. Warm and steady, for when the voice should not be the thing people notice.

Authority and explainers

Charon, Orus, Enceladus, Gacrux. Deep and professional, for product explainers and corporate video.

Energy and social

Puck, Despina, Sadachbia, Sulafat. Playful and upbeat, for short-form content where a flat read loses people.

Storytelling

Algenib, Laomedeia, Erinome, Rasalgethi. Easy to follow over long stretches, for audiobooks and series.

Drama and character

Iapetus, Schedar, Umbriel, Zubenelgenubi. Grand, regal, mysterious and distinctive, for trailers and games.

The rest split between bright and clear (Aoede, Alnilam, Achernar, Algieba, Pulcherrima) and soft and reassuring (Leda, Autonoe, Vindemiatrix, Callirrhoe, Fenrir). Fourteen of the thirty are female voices and sixteen are male.

Audition three from the group that fits, using two sentences of your real script rather than a sample line, then commit before generating the rest. Switching voice halfway through a project means regenerating everything, because listeners hear the seam immediately.

Language coverage

87 languages and regional variants are supported, and 24 of those are fully released: Arabic, Bengali, Dutch, English for both India and the United States, French, German, Hindi, Indonesian, Italian, Japanese, Korean, Marathi, Polish, Portuguese, Romanian, Russian, Spanish, Tamil, Telugu, Thai, Turkish, Ukrainian and Vietnamese.

The other 63 are still being refined, and that group covers British and Australian English, Canadian French, Mexican Spanish, European Portuguese and Mandarin Chinese, alongside a long tail running from Afrikaans to Urdu. They work today, though the output can shift as Google improves them, so build a campaign on the fully released list wherever you can.

Two-speaker dialogue

Label each speaker in the script, assign a voice to each name, and one generation produces the whole exchange. Writing “Sam:” and “Bob:” at the start of each line, then pointing Sam at Kore and Bob at Charon, gives you a two-person conversation in a single pass.

Generating each speaker separately and stitching the takes together afterwards loses the thing that makes a conversation sound real. Speakers in the same generation share context, so the timing, the reactions and the handoffs all land naturally. Keep it in one pass.

How long a script can be

The script box holds up to 5,000 characters, which is roughly five minutes of speech. That covers most voiceover work: a product explainer, a reel narration, a short episode. Longer projects are better split into separate generations anyway, since a single unbroken read gives you nothing to edit around if one line needs redoing.

When you do split a script, break at paragraph ends rather than mid-sentence, and keep every part on the same voice. A fragment that starts mid-clause has no context to set its tone, and a voice change between parts is audible immediately.

Generating a voiceover from here takes a script, an instruction and a voice.

How to use Gemini 2.5 Flash TTS in Picsart

1. Open the AI Playground

Go to the Picsart AI Playground and switch to Audio mode. Everything happens in the browser, so there is nothing to install and no API key to set up.

2. Choose Gemini 2.5 Flash TTS

Select it from the model picker. Auto Mode will otherwise choose a model for you, which may not be the one you want.

3. Paste your script

Add the words you want spoken, up to 5,000 characters. That works out to roughly five minutes of speech, so split anything longer into separate generations.

4. Describe the delivery you want

Lead with your direction, then follow it with the script. Name the tone, the pace and any accent you need, the way you would brief a voice actor. Warm and unhurried, brisk and professional, gentle like a bedtime story: plain language is all it takes.

5. Pick a voice

Choose one of the 30 voices from the dropdown. Audition two or three using a couple of sentences of your real script rather than a sample line, because a voice that stumbles on a product name will keep stumbling on it.

6. Generate and save

Select Generate to see the credit cost, then save the finished audio to a project board so you can reuse it across designs and videos.

Get answers to common questions

Gemini 2.5 Flash TTS is Google’s text-to-speech model for generating natural spoken audio from writing. It offers 30 voices across 87 languages, reads in one voice or two, and takes direction on tone, pace and accent in plain language.

Start generating voiceovers

Gemini 2.5 Flash TTS earns its place for the everyday work: narration for an explainer, a voiceover for a reel, a two-person script for a product walkthrough. Once the audio is generated, drop it straight into the video editor to lay it over your footage. Open the Picsart AI Playground, switch to Audio mode, write your direction ahead of your script, and pick a voice that suits the job.