{"id":260790,"date":"2026-07-30T16:18:14","date_gmt":"2026-07-30T23:18:14","guid":{"rendered":"https:\/\/picsart.com\/blog\/?p=260790"},"modified":"2026-08-14T15:59:59","modified_gmt":"2026-08-14T22:59:59","slug":"what-is-gemini-2-5-flash-tts","status":"publish","type":"post","link":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/","title":{"rendered":"What is Gemini 2.5 Flash TTS and how to use it in Picsart"},"content":{"rendered":"<p>Gemini 2.5 Flash TTS is Google&#8217;s text-to-speech model for turning writing into natural spoken audio, with 30 voices and support for 87 languages. It reads a script in one voice, or performs a conversation between two, and you direct the delivery by describing what you want in ordinary words.<\/p>\n<p>That last part is the thing worth understanding. There are no sliders to balance and no markup language to learn. You write an instruction, the way you would brief a voice actor, and the model performs it.<\/p>\n<h2><span id=\"Gemini_25_Flash_TTS_at_a_glance\">Gemini 2.5 Flash TTS at a glance<\/span><\/h2>\n<figure class=\"wp-block-table\">\n<table style=\"border-collapse: collapse; width: 100%; table-layout: auto;\">\n<tbody>\n<tr>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #000000; font-weight: bold; white-space: nowrap;\">Model ID<\/td>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #141414;\">gemini-2.5-flash-tts<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #000000; font-weight: bold; white-space: nowrap;\">Best for<\/td>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #141414;\">Fast, low-cost everyday voiceovers<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #000000; font-weight: bold; white-space: nowrap;\">Input and output<\/td>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #141414;\">Text in, audio out<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #000000; font-weight: bold; white-space: nowrap;\">Voices<\/td>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #141414;\">30<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #000000; font-weight: bold; white-space: nowrap;\">Languages<\/td>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #141414;\">87, including 24 fully released<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #000000; font-weight: bold; white-space: nowrap;\">Speakers<\/td>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #141414;\">One voice, or two for a conversation<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #000000; font-weight: bold; white-space: nowrap;\">Audio formats<\/td>\n<td style=\"border: 1px solid #333333; padding: 10px 14px; text-align: left; vertical-align: top; color: #ffffff; background: #141414;\">MP3, WAV, OGG and PCM<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<h2><span id=\"Tell_it_how_to_read_not_just_what_to_read\">Tell it how to read, not just what to read<\/span><\/h2>\n<p>Two separate things go in. The script holds the words you want spoken. The instruction describes how to say them, and it never gets read aloud. That split is why the model takes direction so well, because your stage notes stay out of the script and the voice cannot read your instructions back to you.<\/p>\n<p>Writing &#8220;say this in a curious way&#8221; ahead of your script is enough to change the whole read. Warm and unhurried, brisk and professional, gentle like a bedtime story: describe the performance and the model works toward it. This is also where accents come from, so asking for a regional accent in the instruction is what produces one. Changing the language setting on its own will not.<\/p>\n<h2><span id=\"Directing_pace_tone_and_emotion\">Directing pace, tone and emotion<\/span><\/h2>\n<p>For a change partway through a line, drop a cue in square brackets exactly where the shift should happen. Writing <strong>[extremely fast]<\/strong> in front of a legal disclaimer races through it the way a real voice actor would. <strong>[whispers]<\/strong> drops the delivery to a hush mid-sentence.<\/p>\n<p>Three kinds of direction do most of the work:<\/p>\n<ul>\n<li><strong>Tone and style.<\/strong> Ask for a mood, a register or a whisper, and the delivery follows.<\/li>\n<li><strong>Pace.<\/strong> Speeding up or slowing down also sharpens pronunciation, which helps with tricky names and numbers.<\/li>\n<li><strong>Accent.<\/strong> Name the accent you want in the instruction rather than relying on the language setting.<\/li>\n<\/ul>\n<p>The model handles poetry, news copy and storytelling convincingly, and performs a specific emotion when you name one.<\/p>\n<h2><span id=\"Language_coverage\">Language coverage<\/span><\/h2>\n<p>87 languages and regional variants are supported, and 24 of those are fully released: Arabic, Bengali, Dutch, English for both India and the United States, French, German, Hindi, Indonesian, Italian, Japanese, Korean, Marathi, Polish, Portuguese, Romanian, Russian, Spanish, Tamil, Telugu, Thai, Turkish, Ukrainian and Vietnamese.<\/p>\n<p>The other 63 are still being refined, and that group covers British and Australian English, Canadian French, Mexican Spanish, European Portuguese and Mandarin Chinese, alongside a long tail running from Afrikaans to Urdu. They work today, though the output can shift as Google improves them, so build a campaign on the fully released list wherever you can.<\/p>\n<h2><span id=\"Two-speaker_dialogue\">Two-speaker dialogue<\/span><\/h2>\n<p>Label each speaker in the script, assign a voice to each name, and one generation produces the whole exchange. Writing &#8220;Sam:&#8221; and &#8220;Bob:&#8221; at the start of each line, then pointing Sam at Kore and Bob at Charon, gives you a two-person conversation in a single pass.<\/p>\n<p>Generating each speaker separately and stitching the takes together afterwards loses the thing that makes a conversation sound real. Speakers in the same generation share context, so the timing, the reactions and the handoffs all land naturally. Keep it in one pass.<\/p>\n<h2><span id=\"How_long_a_script_can_be\">How long a script can be<\/span><\/h2>\n<p>The script box holds up to 5,000 characters, which is roughly five minutes of speech. That covers most voiceover work: a product explainer, a reel narration, a short episode. Longer projects are better split into separate generations anyway, since a single unbroken read gives you nothing to edit around if one line needs redoing.<\/p>\n<p>When you do split a script, break at paragraph ends rather than mid-sentence, and keep every part on the same voice. A fragment that starts mid-clause has no context to set its tone, and a voice change between parts is audible immediately.<\/p>\n<p>Generating a voiceover from here takes a script, an instruction and a voice.<\/p>\n<section class=\"section_faq\" id=\"faq-faq-6a8d3ff97df95\">\n            <h2 class=\"faq_title\" id=\"Get_answers_to_common_questions\">Get answers to common questions<\/h2>\n    \n    <div class=\"faq_items\">\n                    <div class=\"faq_item faq_item--active\">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"true\">\n                    <span class=\"faq_question_text\">What is Gemini 2.5 Flash TTS?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"false\">\n                    <div class=\"faq_answer_content\"><p>Gemini 2.5 Flash TTS is Google&#8217;s text-to-speech model for generating natural spoken audio from writing. It offers 30 voices across 87 languages, reads in one voice or two, and takes direction on tone, pace and accent in plain language.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">How many voices does Gemini 2.5 Flash TTS have?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>30 voices, made up of 14 female and 16 male. Gemini 2.5 Pro TTS offers the same set, so switching between the two models will not change the voice you picked.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">How many languages does Gemini 2.5 Flash TTS support?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>87 languages and regional variants. 24 of those are fully released, including English, Spanish, French, German, Hindi, Japanese, Korean and Portuguese. The remaining 63 are still being refined, so their output can change over time.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">Can Gemini 2.5 Flash TTS generate a conversation between two speakers?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>Yes. Label each speaker in the script, assign a voice to each name, and the model produces the full exchange in a single generation. Keeping it in one pass is what makes the timing sound natural.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">What is the difference between Gemini 2.5 Flash TTS and Pro TTS?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>Both carry the same 30 voices and both handle two speakers. Flash is built for speed and everyday work at low cost. Pro is built for tighter control on podcasts, audiobooks and customer support.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">Can Gemini 2.5 Flash TTS clone my voice?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>No. It works from a set of 30 prebuilt voices and does not accept a reference recording. Pick the prebuilt voice closest to the read you want, then shape the delivery with your written instruction.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n            <\/div>\n<\/section>\n\n<script type=\"application\/ld+json\">\n{\n    \"@context\": \"https:\/\/schema.org\",\n    \"@type\": \"FAQPage\",\n    \"mainEntity\": [\n        {\n            \"@type\": \"Question\",\n            \"name\": \"What is Gemini 2.5 Flash TTS?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Gemini 2.5 Flash TTS is Google&#8217;s text-to-speech model for generating natural spoken audio from writing. It offers 30 voices across 87 languages, reads in one voice or two, and takes direction on tone, pace and accent in plain language.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"How many voices does Gemini 2.5 Flash TTS have?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"30 voices, made up of 14 female and 16 male. Gemini 2.5 Pro TTS offers the same set, so switching between the two models will not change the voice you picked.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"How many languages does Gemini 2.5 Flash TTS support?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"87 languages and regional variants. 24 of those are fully released, including English, Spanish, French, German, Hindi, Japanese, Korean and Portuguese. The remaining 63 are still being refined, so their output can change over time.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Can Gemini 2.5 Flash TTS generate a conversation between two speakers?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Yes. Label each speaker in the script, assign a voice to each name, and the model produces the full exchange in a single generation. Keeping it in one pass is what makes the timing sound natural.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"What is the difference between Gemini 2.5 Flash TTS and Pro TTS?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Both carry the same 30 voices and both handle two speakers. Flash is built for speed and everyday work at low cost. Pro is built for tighter control on podcasts, audiobooks and customer support.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Can Gemini 2.5 Flash TTS clone my voice?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"No. It works from a set of 30 prebuilt voices and does not accept a reference recording. Pick the prebuilt voice closest to the read you want, then shape the delivery with your written instruction.\"\n            }\n        }\n    ]\n}<\/script>\n\n<script>\n(function() {\n    var container = document.getElementById('faq-faq-6a8d3ff97df95');\n    if (!container) return;\n\n    var items = container.querySelectorAll('.faq_item');\n    items.forEach(function(item) {\n        var button = item.querySelector('.faq_question');\n        var answer = item.querySelector('.faq_answer');\n        if (!button || !answer) return;\n\n        button.addEventListener('click', function() {\n            var isActive = item.classList.contains('faq_item--active');\n\n            if (isActive) {\n                item.classList.remove('faq_item--active');\n                button.setAttribute('aria-expanded', 'false');\n                answer.setAttribute('aria-hidden', 'true');\n                answer.setAttribute('data-collapsed', '');\n            } else {\n                items.forEach(function(other) {\n                    var otherBtn = other.querySelector('.faq_question');\n                    var otherAnswer = other.querySelector('.faq_answer');\n                    other.classList.remove('faq_item--active');\n                    if (otherBtn) otherBtn.setAttribute('aria-expanded', 'false');\n                    if (otherAnswer) {\n                        otherAnswer.setAttribute('aria-hidden', 'true');\n                        otherAnswer.setAttribute('data-collapsed', '');\n                    }\n                });\n                item.classList.add('faq_item--active');\n                button.setAttribute('aria-expanded', 'true');\n                answer.removeAttribute('data-collapsed');\n                answer.setAttribute('aria-hidden', 'false');\n            }\n        });\n    });\n})();\n<\/script>\n\n<h2><span id=\"Start_generating_voiceovers\">Start generating voiceovers<\/span><\/h2>\n<p>Gemini 2.5 Flash TTS earns its place for the everyday work: narration for an explainer, a voiceover for a reel, a two-person script for a product walkthrough. Open the <a href=\"https:\/\/picsart.com\/ai-playground\/\">Picsart AI Playground<\/a>, switch to Audio mode, write your direction ahead of your script, and pick a voice that suits the job.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Gemini 2.5 Flash TTS is Google&#8217;s text-to-speech model for turning writing into natural spoken audio, with 30 voices and support for 87 languages. It reads a script in one voice, or performs a conversation between two, and you direct the delivery by describing what you want in ordinary words. That last part is the thing &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;What is Gemini 2.5 Flash TTS and how to use it in Picsart&#8221;<\/span><\/a><\/p>\n","protected":false},"author":146,"featured_media":244536,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_yoast_wpseo_title":"What is Gemini 2.5 Flash TTS","_yoast_wpseo_metadesc":"Gemini 2.5 Flash TTS explained: 30 voices, 87 languages and two-speaker dialogue. Learn how to direct tone, pace and accent, and how to use it in Picsart.","faq_show":true,"faq_enable_schema":true,"how_to_show":false,"how_to_show_on_single":false,"how_to_enable_schema":false,"how_to_is_upload":true,"faq_title":"Get answers to common questions","how_to_title":"","how_to_layout":"default","how_to_cta_text":"","how_to_cta_url":"","how_to_image_alt":"","how_to_display_image":0,"faq_items":null,"how_to_steps":[],"prompt_box_show":false,"prompt_box_placeholder":"","prompt_box_deeplink":"","prompt_box_submit_label":"","footnotes":""},"categories":[3181,1669],"tags":[3695],"class_list":["post-260790","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai","category-inspiration","tag-audio-generation","entry"],"acf":{"footer_banner_name":"Start your design in Picsart","footer_banner_link_":"\/","footer_banner_button_text_":"Get Started","faq_show":true,"faq_title":"Get answers to common questions","faq_enable_schema":true,"faq_items":[{"question":"What is Gemini 2.5 Flash TTS?","answer":"Gemini 2.5 Flash TTS is Google's text-to-speech model for generating natural spoken audio from writing. It offers 30 voices across 87 languages, reads in one voice or two, and takes direction on tone, pace and accent in plain language."},{"question":"How many voices does Gemini 2.5 Flash TTS have?","answer":"30 voices, made up of 14 female and 16 male. Gemini 2.5 Pro TTS offers the same set, so switching between the two models will not change the voice you picked."},{"question":"How many languages does Gemini 2.5 Flash TTS support?","answer":"87 languages and regional variants. 24 of those are fully released, including English, Spanish, French, German, Hindi, Japanese, Korean and Portuguese. The remaining 63 are still being refined, so their output can change over time."},{"question":"Can Gemini 2.5 Flash TTS generate a conversation between two speakers?","answer":"Yes. Label each speaker in the script, assign a voice to each name, and the model produces the full exchange in a single generation. Keeping it in one pass is what makes the timing sound natural."},{"question":"What is the difference between Gemini 2.5 Flash TTS and Pro TTS?","answer":"Both carry the same 30 voices and both handle two speakers. Flash is built for speed and everyday work at low cost. Pro is built for tighter control on podcasts, audiobooks and customer support."},{"question":"Can Gemini 2.5 Flash TTS clone my voice?","answer":"No. It works from a set of 30 prebuilt voices and does not accept a reference recording. Pick the prebuilt voice closest to the read you want, then shape the delivery with your written instruction."}],"how_to_show":false,"how_to_show_on_single":false,"how_to_title":"","how_to_layout":"default","how_to_steps":null,"how_to_enable_schema":false,"how_to_is_upload":true,"how_to_cta_text":"","how_to_cta_url":"","how_to_display_image":"","how_to_image_alt":"","prompt_box_show":false,"prompt_box_placeholder":"","prompt_box_deeplink":"https:\/\/picsart.com\/create\/editor?category=miniapps&app=com.picsart.aura","prompt_box_submit_label":"Create","try_prompt_show":false,"try_prompt_title":"Try this prompt","try_prompt_text":"","try_prompt_deeplink":"","tips_show":false,"tips_title":"Tips for best results","tips_items":null,"cta_banner_show":false,"cta_banner_title":"Need more space?","cta_banner_subtitle":"Extend any image in any direction with AI.","cta_banner_button_label":"Expand image","cta_banner_button_url":"","related_tools_title":"Related tools","related_tools_items":null,"post_level":""},"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v25.5 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>What is Gemini 2.5 Flash TTS<\/title>\n<meta name=\"description\" content=\"Gemini 2.5 Flash TTS explained: 30 voices, 87 languages and two-speaker dialogue. Learn how to direct tone, pace and accent, and how to use it in Picsart.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"What is Gemini 2.5 Flash TTS\" \/>\n<meta property=\"og:description\" content=\"Gemini 2.5 Flash TTS explained: 30 voices, 87 languages and two-speaker dialogue. Learn how to direct tone, pace and accent, and how to use it in Picsart.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/\" \/>\n<meta property=\"og:site_name\" content=\"Picsart Blog\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/picsart\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-30T23:18:14+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-14T22:59:59+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/cdnblog.picsart.com\/2026\/03\/CR6776.-How-to-Make-an-AI-Voice-with-Picsart-_-1200_800.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"800\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Julia Tovmasyan\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@PicsArtStudio\" \/>\n<meta name=\"twitter:site\" content=\"@PicsArtStudio\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Julia Tovmasyan\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"4 minutes\" \/>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"What is Gemini 2.5 Flash TTS","description":"Gemini 2.5 Flash TTS explained: 30 voices, 87 languages and two-speaker dialogue. Learn how to direct tone, pace and accent, and how to use it in Picsart.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/","og_locale":"en_US","og_type":"article","og_title":"What is Gemini 2.5 Flash TTS","og_description":"Gemini 2.5 Flash TTS explained: 30 voices, 87 languages and two-speaker dialogue. Learn how to direct tone, pace and accent, and how to use it in Picsart.","og_url":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/","og_site_name":"Picsart Blog","article_publisher":"https:\/\/www.facebook.com\/picsart","article_published_time":"2026-07-30T23:18:14+00:00","article_modified_time":"2026-08-14T22:59:59+00:00","og_image":[{"width":1200,"height":800,"url":"https:\/\/cdnblog.picsart.com\/2026\/03\/CR6776.-How-to-Make-an-AI-Voice-with-Picsart-_-1200_800.png","type":"image\/png"}],"author":"Julia Tovmasyan","twitter_card":"summary_large_image","twitter_creator":"@PicsArtStudio","twitter_site":"@PicsArtStudio","twitter_misc":{"Written by":"Julia Tovmasyan","Est. reading time":"4 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/#article","isPartOf":{"@id":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/"},"author":{"name":"Julia Tovmasyan","@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/person\/74b70f3125250c23596a5306775b702d"},"headline":"What is Gemini 2.5 Flash TTS and how to use it in Picsart","datePublished":"2026-07-30T23:18:14+00:00","dateModified":"2026-08-14T22:59:59+00:00","mainEntityOfPage":{"@id":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/"},"wordCount":758,"publisher":{"@id":"https:\/\/picsart.com\/blog\/ko\/#organization"},"image":{"@id":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/#primaryimage"},"thumbnailUrl":"https:\/\/cdnblog.picsart.com\/2026\/03\/CR6776.-How-to-Make-an-AI-Voice-with-Picsart-_-1200_800.png","keywords":["Audio Generation"],"articleSection":["AI","Inspirational"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/","url":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/","name":"What is Gemini 2.5 Flash TTS","isPartOf":{"@id":"https:\/\/picsart.com\/blog\/ko\/#website"},"primaryImageOfPage":{"@id":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/#primaryimage"},"image":{"@id":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/#primaryimage"},"thumbnailUrl":"https:\/\/cdnblog.picsart.com\/2026\/03\/CR6776.-How-to-Make-an-AI-Voice-with-Picsart-_-1200_800.png","datePublished":"2026-07-30T23:18:14+00:00","dateModified":"2026-08-14T22:59:59+00:00","description":"Gemini 2.5 Flash TTS explained: 30 voices, 87 languages and two-speaker dialogue. Learn how to direct tone, pace and accent, and how to use it in Picsart.","breadcrumb":{"@id":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/#primaryimage","url":"https:\/\/cdnblog.picsart.com\/2026\/03\/CR6776.-How-to-Make-an-AI-Voice-with-Picsart-_-1200_800.png","contentUrl":"https:\/\/cdnblog.picsart.com\/2026\/03\/CR6776.-How-to-Make-an-AI-Voice-with-Picsart-_-1200_800.png","width":1200,"height":800,"caption":"ai voice picsart"},{"@type":"BreadcrumbList","@id":"https:\/\/picsart.com\/blog\/what-is-gemini-2-5-flash-tts\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/picsart.com\/blog\/"},{"@type":"ListItem","position":2,"name":"What is Gemini 2.5 Flash TTS and how to use it in Picsart"}]},{"@type":"WebSite","@id":"https:\/\/picsart.com\/blog\/ko\/#website","url":"https:\/\/picsart.com\/blog\/ko\/","name":"Picsart Blog","description":"Keep up with the latest news in photo editing, digital photography, and art trends.","publisher":{"@id":"https:\/\/picsart.com\/blog\/ko\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/picsart.com\/blog\/ko\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/picsart.com\/blog\/ko\/#organization","name":"PicsArt Inc.","url":"https:\/\/picsart.com\/blog\/ko\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/logo\/image\/","url":"https:\/\/cdnblog.picsart.com\/2016\/02\/PicsArt-logo.png","contentUrl":"https:\/\/cdnblog.picsart.com\/2016\/02\/PicsArt-logo.png","width":195,"height":43,"caption":"PicsArt Inc."},"image":{"@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/picsart","https:\/\/x.com\/PicsArtStudio","https:\/\/www.instagram.com\/picsart","https:\/\/www.linkedin.com\/company\/picsart-photo-studio","https:\/\/www.pinterest.com\/picsart"]},{"@type":"Person","@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/person\/74b70f3125250c23596a5306775b702d","name":"Julia Tovmasyan","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/person\/image\/","url":"https:\/\/cdnblog.picsart.com\/2026\/03\/3285C16C-FD87-4868-A2F0-04B6A0815CE1-150x150.jpg","contentUrl":"https:\/\/cdnblog.picsart.com\/2026\/03\/3285C16C-FD87-4868-A2F0-04B6A0815CE1-150x150.jpg","caption":"Julia Tovmasyan"}}]}},"featured_image":{"url":"https:\/\/cdnblog.picsart.com\/2026\/03\/CR6776.-How-to-Make-an-AI-Voice-with-Picsart-_-1200_800.png","dimensions":[]},"_links":{"self":[{"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/posts\/260790","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/users\/146"}],"replies":[{"embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/comments?post=260790"}],"version-history":[{"count":10,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/posts\/260790\/revisions"}],"predecessor-version":[{"id":262654,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/posts\/260790\/revisions\/262654"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/media\/244536"}],"wp:attachment":[{"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/media?parent=260790"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/categories?post=260790"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/tags?post=260790"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}