Google released two new text to speech models on September 23, 2026, and they sit far closer together than the names suggest. Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS build voices the same way, take direction the same way, and record for the same length. Moving a script from one to the other changes nothing about how you work.
That leaves two differences that decide anything at all. Flash TTS speaks 130 languages against Flash-Lite’s 101. And Flash-Lite costs less to run.
So the useful question is not which model is better. It is whether your script needs one of the 29 languages the flagship covers alone, or the extra polish it brings to harder jobs like layered dialogue and regional accents. If it does not, the cheaper model gives you the same result.
The differences at a glance
| Gemini 3.8 Flash TTS | Gemini 3.8 Flash-Lite TTS | |
|---|---|---|
| Languages | 130 | 101 |
| Overall quality ranking | 1st | 2nd |
| Custom voice ranking | 1st | Not ranked |
| Overlapping dialogue | Works best here | Supported |
| Built for | Rich performance, accents, long narration | Speed, volume, everyday voiceover |
Those rows are the whole list. Everything else about the two models matches, which is why the next section is the longer one.
What both models do exactly the same
This list is longer than the list of differences, and it is the reason the choice is lower stakes than it looks.
Directing the delivery. On both models you set a mood for a whole line, like whispered urgently or warm and enthusiastic, and you place small sounds exactly where you want them, like a laugh, a sigh, a breath or a pause. Your notes never get read aloud. A script written for one model works unchanged on the other.
Every voice option. Both reach the 30 ready-made voices, hundreds more in the extended library, custom voices you create by describing them in words, and voice replication that rebuilds a speaker’s voice from a short sample with consent checks. Google’s own capability list marks all of it available on both.
Two-speaker scenes. Both stage a conversation between two voices from a single script, with natural turn taking, so a podcast-style exchange does not have to be recorded twice and stitched together.
Recording length. Both produce the same amount of audio in one run, so neither gives you more room than the other.
The practical things. File formats and bulk processing options match as well.
Where they actually differ
Language coverage, by 29. Flash TTS reads 130 languages and Flash-Lite reads 101. This is the largest gap between them and the one most likely to rule a model out before anything else does.
Cost. Flash-Lite generates audio for about a third less than the flagship. Both rates rise at the start of 2027, so the gap between them stays the same either side of that date.
Polish on demanding material. Google points Flash TTS at audiobooks, studio narration, complex multi-speaker scenes, heavy character acting, tricky pronunciation and regional accents. Flash-Lite is aimed at bulk work, voice assistants, read-aloud features and everyday single-voice recordings.
Overlapping dialogue. Both handle listener reactions layered inside a speaker’s line. Google notes that genuinely simultaneous or interrupted speech works best on the flagship, which is a stated preference rather than a hard limit.
The 29 languages Flash-Lite does not cover
Google publishes a language by language breakdown, and the gap in it is worth reading before you pick on price. These 29 run on Flash TTS only:
Banjar in Arabic script, Bashkir, Bemba, Burmese, Crimean Tatar, Dyula, Dzongkha, Finnish, Guarani, Igbo, Kabyle, Latgalian, Lithuanian, Luxembourgish, Minangkabau in Latin script, Occitan, Pangasinan, Sindhi, Slovenian, Somali, Southern Sotho, Swahili, Swati, Swedish, Tajik, Thai, Tigrinya, Tosk Albanian and Uyghur.
Several of those are large markets rather than edge cases. Thai, Swedish, Finnish, Lithuanian and Slovenian all sit on the flagship-only side, as do Swahili and Somali.
There is an awkward consequence worth naming. Flash-Lite is the model built for high-volume dubbing and localization, and it is also the model with the shorter language list. If you are localizing a catalog, check your target languages before you build around the cheaper option, because the saving disappears the moment you need a second model to fill the gaps.
How long a single recording can be
Google describes long narration running for hours with the voice holding steady, and the steadiness is real. The length is worth reading carefully, though.
One recording run produces roughly 11 minutes of speech. An audiobook or a feature-length narration is therefore made of pieces on either model rather than one continuous take.
What the models genuinely give you is consistency across those pieces. The voice, its tone, its volume and even the sense of the room stay the same from one segment to the next, which is the part that used to break. Plan in 11-minute chunks and the promise holds.
Which one fits your work
Choose Flash-Lite TTS for volume in a widely spoken language. Voice assistants, read-aloud features, bulk narration and everyday single-voice work all sit squarely in what it was built for. You keep custom voices and voice replication, so nothing about your voice choices has to change.
Choose Flash TTS when the material is demanding or the language is on the exclusive list. Layered dialogue, heavy character acting, difficult pronunciation, regional accents, and long narration where the voice has to hold are what the higher cost buys you.
Try both if you are unsure, because it is unusually easy here. The same script and the same voice work on either model, so comparing them costs you a setting rather than an afternoon. Record the same 30 seconds on each, listen to them back to back, and let your own material decide.
Get answers to common questions
Yes. Both models create custom voices from a written description and rebuild a speaker’s voice from a short sample with consent checks. Flash-Lite gives up nothing on voices compared with the flagship.
Hear what the Gemini voices sound like
Google’s newest speech models have not reached Picsart yet, but the Gemini text to speech family has. The AI Playground holds 198 models from 34 providers behind a single prompt bar, Gemini 2.5 Pro TTS among them, so you can hear how these voices read your own script today. The model catalog shows what each one is built for.
Start with the script you are actually stuck on. It will tell you more in five minutes than any comparison table.