Contents
ElevenLabs text to speech turns a written script into studio-quality spoken audio, and the model you pick decides what kind of performance you get. Eleven v3 acts a script with emotion and direction. Eleven Multilingual v2 narrates it cleanly and holds steady across long recordings. Both cost the same, so the choice is about the read rather than the budget.
This guide covers what each model is built for, how to direct the delivery with audio tags, how to generate a two-person conversation, and how to run the whole thing in a browser.
The two models and what each one is built for
Eleven v3 is the expressive one. It covers more than 70 languages, carries audio tags and multi-speaker dialogue, takes 5,000 characters per generation for about five minutes of speech, and is still marked experimental.
Eleven Multilingual v2 is the stable, professional option. It covers 29 languages, takes 10,000 characters for about ten minutes in one pass, and is what ElevenLabs themselves recommend for content creation, audiobooks and video narration.
Both run at 3 credits per 1,000 characters, so the choice is about the read rather than the cost.
That surprises people. Eleven v3 is newer and more capable, but for a straight voiceover on a finished video, consistency across a long read matters more than dramatic range, and Multilingual v2 also gives you double the script length in one pass.
The model reads the room, not just the words
Before you write a single instruction, the voice is already responding to what your script says. ElevenLabs voices pick up emotional cues in the text and adapt delivery to both the sentence at hand and the wider context around it, which is what stops a line landing in the wrong tone when the subject changes.
That matters in practice, because a well-written script often needs no direction at all. Tags are for the moments where you want something the words alone would not produce.
Audio tags are what set Eleven v3 apart. Write a cue in square brackets, drop it inline in the script, and the delivery changes at exactly that point.
[slowly] Back then... [chuckles] we had no phones.
[whispers] Just dirt roads and [coughs] big dreams. [sad] Then it happened.
The tags group into three kinds:
- Emotion. [sad], [angry], [curious], [sarcastically], [mischievously]
- Delivery and pace. [whispers], [shouts], [slowly]
- Human reactions. [laughs], [chuckles], [giggles], [sighs], [clears throat], [coughs]
Capital letters carry weight too, so a shouted line reads as And GOOOOAL! rather than leaning on a tag alone. Character colour comes from stacked cues like [laughs wickedly] and [evil laugh].
Longer tag lists circulate online, but they are community-collected and not all of them fire. Sticking to the documented set is the difference between a script that performs and one that reads your notes aloud.
Writing a script the model reads well
Punctuation is performance direction. Commas and full stops create the pauses, so a run-on sentence gets read as a run-on sentence. Spell out anything unusual the way you want it said, including acronyms, product names and numbers, and respell a stubborn word phonetically if it keeps coming out wrong.
Give the model enough runway as well. A clipped fragment offers no context for setting tone, so a few full sentences almost always read better than one short line.
Generating a two-person conversation
Eleven v3 supports natural multi-speaker dialogue. Label each speaker on its own line and let a single generation carry the whole exchange:
Mark
Hey Chris... Knock knock.
Chris
[chuckles] I'm not doing this AGAIN!
Mark
[laughing] Come on, PLEASE! I promise you'll love this one.
Speakers in one generation share context, so timing, interruptions and reactions all land naturally, and even overlapping speech survives. Generating each part separately and stitching the takes together loses precisely that, and the joins are audible.
Choosing a voice
Voices are grouped by the job they are built for, which is a faster way in than scrolling a list. Narration voices carry audiobooks and podcasts. Conversational voices suit informal, spoken-word scenarios. Character voices are built for cartoons and games. Social media voices are made to hold attention in short-form video, and advertisement voices aim at recall and action.
What the finished audio sounds like
Speech exports as MP3 at 44.1 kHz and 128 kbps, which is the same quality you would expect from a music file and more than enough for social video, podcasts and voiceover work. Once it is generated, drop it into the video editor to lay it over your footage.
When output goes wrong
| Problem | Fix |
|---|---|
| Audio tags ignored | Switch to Eleven v3 |
| Word mispronounced | Respell it phonetically |
| Read is too flat | Add a delivery tag, give the line more context |
| Voice shifts between parts | Same model and voice throughout, split at paragraph ends |
| Quality drifts on long audio | Move to Multilingual v2 |
The rest of the ElevenLabs audio models
Text to speech is one part of a wider set. Eleven Dubbing is the standout, because it is free and it takes a finished video into another language with the voices matched. Eleven STS v2 swaps one voice for another while keeping the original timing, and Eleven Audio Isolation strips background noise from a messy recording. For audio that does not exist yet, ElevenLabs SFX v2 makes sound effects up to 15 seconds long, the AI music generator runs Music v2 for full tracks, and Eleven Voice Design v3 builds a brand new voice from a description.
Generating a voiceover takes a script, a model and a voice.
How to use ElevenLabs text to speech with Picsart
1. Open the AI Playground
Go to the Picsart AI Playground and switch to Audio mode. Everything runs in the browser, so there is nothing to install and nothing to set up.
2. Choose your model
Pick Eleven v3 for an expressive, performed read, or Eleven Multilingual v2 for steady narration and a longer script in one pass. Auto Mode will otherwise choose a model for you.
3. Paste your script
Add the words you want spoken. Punctuate deliberately, because commas and full stops create the pauses, and spell out acronyms and numbers the way you want them said. On Eleven v3 you can write audio tags inline in square brackets.
4. Set the language and accent
Both are plain text fields rather than menus, so describe what you need. Naming a regional accent here is what produces one.
5. Pick a voice
Audition two or three voices on a couple of sentences of your real script rather than on a sample line. A voice that fumbles a product name will keep fumbling it, and that only shows up on your own copy.
6. Generate and save
Select Generate to see the credit cost, then save the audio to a project board so you can reuse it across designs and videos.
Get answers to common questions
Eleven v3 for storytelling, character work and anything that needs a performed read, since it is the model that responds to audio tags. Eleven Multilingual v2 for voiceovers, audiobooks and long recordings that need a consistent voice. Both cost the same in Picsart, so the choice is about the read rather than the budget.
Start generating voiceovers
ElevenLabs text to speech earns its place on the everyday work: narration for an explainer, a voiceover for a reel, a two-person script for a product walkthrough. Open the Picsart AI Playground, pick the model that matches the read, and write your script so the punctuation does half the directing for you.