ElevenLabs v4 and Eleven v4 Turbo are one voice engine at two speeds.

The choice comes down to one question: does your audio need to answer back in real time?

  • Eleven v4 Turbo: live conversation and agents.
  • Eleven v4: produced audio like audiobooks, character work, voiceovers and dubbing.

Everything else matches on Picsart, from audio tags to languages and settings.

What Eleven v4 and Eleven v4 Turbo share

On Picsart, the settings are identical. Both models take a prompt of up to 10,000 characters, a voice from the same catalog, an optional language, stability and similarity sliders from 0 to 1 in steps of 0.05, and an optional Character Timings switch that is off by default.

Which model fits your project

The fastest way to choose is to start from the job. Anything that has to answer a person live goes to Turbo, and anything you produce and publish goes to Eleven v4.

Project Pick Why
Audiobook or long narration Eleven v4 Steady pacing across long scripts
Voiceover or ad read Eleven v4 No vocal drift on redone lines
Customer support agent Eleven v4 Turbo ~150 ms to first speech
Agent reading live text Eleven v4 Turbo Audio starts before the sentence ends

Directing a performance with ElevenLabs v4 audio tags

This is where ElevenLabs v4 feels new, and it is fully shared. The engine is built to read a script the way a voice actor would. It tracks who is speaking, what just happened and how each line should land, then delivers it dramatic, tender, urgent, comedic or conversational without losing the speaker’s identity.

You direct it inside the script itself. Inline tags like [laughs], [whispers] or [door slams] shape the delivery and drop in sound effects, and natural-language direction works too, such as [said angrily in French accent]. ElevenLabs says v4 follows tag sequences more reliably than v3, sound effects included.

Turbo reads exactly the same tags. It is Eleven v4 tuned for speed, with the same audio tags and languages, so a script you direct for one model carries straight over to the other. Here are two lines to start with on either model.

Upbeat announcement

[excited, happy] Big news, everyone. The summer menu drops this Friday. And yes, the mango sticky rice is finally back.

Sleepy character beat

[sleepy drowsy voice] Mmm, five more minutes. [yawning] I promise I’ll get up right after this dream finishes.

Notice how the tags stack. The first line sets one mood for the whole read, while the second shifts mid-line from a drowsy voice into an actual yawn. Mix scene sounds like [light rain] or [phone buzzing] into a script and the model weaves them into the read.

Long scripts: where Eleven v4 earns its place

Long-form audio is the clearest reason to choose Eleven v4. ElevenLabs built it for context stitching, which keeps pacing and delivery steady across a script of any length. The goal is a full audiobook that sounds like a single take rather than a chain of separate recordings.

That matters on Picsart because each generation caps the prompt at 10,000 characters on both models. A novel, a course or a long documentary script arrives in pieces. ElevenLabs says request stitching, the chaining of generations for long-form work, is significantly more reliable in Eleven v4.

Regeneration is the other half of the story. Redo a line once or fifty times in v4, and it is still the same person speaking, with no vocal drift. Narrators and editors can fix a single fluffed sentence in chapter twelve without the voice suddenly sounding like someone else.

Choose Eleven v4 for

  • Audiobooks and long narration
  • Character performances
  • Voiceovers for videos and ads
  • Dubbing and localized versions of a script

Real-time conversation: what Eleven v4 Turbo is built for

Eleven v4 Turbo exists for a problem every voice agent builder knows. ElevenLabs points out that teams historically had to choose between fast agents and expressive ones. Turbo combines the two, carrying the expressive range of v4 into low-latency work.

The numbers tell the story. Turbo’s median inference latency sits around 100 ms, which ElevenLabs describes as faster than the average pause between two people talking. Median time to first speech is around 150 ms, and the model stays consistent across long interactions.

Turbo also supports bidirectional streaming. Text goes in as a language model writes it, and audio starts coming back before the sentence is even finished. That is the loop conversational agents run on, where every half-second of silence feels like hesitation.

Speed does not flatten the delivery. ElevenLabs says confirmations, escalations and holds land differently from one another instead of reading identically. Its own examples include a warm, reassuring healthcare agent that pronounces medical terms accurately and a fast-talking, slang-wielding gaming agent.

Support agent hold message

[warm, reassuring] Thanks for waiting. I’ve found your booking, and everything looks good on my end. Give me one second to confirm the time.

Languages, accents and voice cloning on both models

Here the two models match again, so this section stays short. Both Eleven v4 and Eleven v4 Turbo speak 90+ languages. A voice recorded in one language can speak any other fluently with a native accent while keeping its identity, and accent adherence is stronger, with no drift back toward the source accent mid-generation.

That makes either model a natural fit for localization, where a brand voice needs to sound right in every supported language. ElevenLabs highlights Turbo speaking Japanese, Spanish and Portuguese fluently with a native accent, which helps agents that serve customers across regions.

Cloning improvements also land on both. Speaker similarity is significantly better, Instant Voice Clones work from 10 seconds of audio, and identity holds across generations, dialogue, narration and regenerated lines. Professional Voice Clones, absent from v3, are back and work identically across both models.

Both models also share higher underlying audio fidelity and cleaner output. On Picsart, they draw on the same voice catalog, so a voice you audition on Eleven v4 is waiting for you on Turbo.

How to try Eleven v4 and Eleven v4 Turbo on Picsart

Both models live in the Picsart AI Playground, and the steps are the same for each. Open Eleven v4 in the AI Playground for produced audio, or Eleven v4 Turbo in the AI Playground for the low-latency version.

1. Open the model

Launch the AI Playground with Eleven v4 or Eleven v4 Turbo already selected in audio mode.

2. Write your script with tags

Paste or type your script, up to 10,000 characters, and add audio tags where you want the delivery to shift.

3. Pick a voice

Choose a voice from the catalog. The same catalog serves both models.

4. Generate

Run the generation and listen back. Swap models and run the same script again to hear the difference for yourself.


Get answers to common questions

They are the same engine tuned for different jobs. Eleven v4 Turbo is built for low latency, with around 100 ms median inference and bidirectional streaming for agents. Eleven v4 is the pick for produced, long-form audio such as audiobooks, voiceovers and dubbing.

Which ElevenLabs v4 model to pick

The decision is refreshingly simple. Choose Eleven v4 Turbo for anything that has to talk back in real time, from support agents to game characters. Choose ElevenLabs v4 for anything you produce, polish and publish, from a single ad read to a full audiobook.

Since the tags, voices and settings match, testing both costs you a click rather than a rewrite. Write one tagged script, run it on each model, and let your ears make the call. Start with Eleven v4 in the AI Playground and give your next script a voice.