Seed Audio 1.0 is ByteDance’s audio creation model, released in July 2026, and it treats a piece of audio as a scene rather than a stack of separate files. Voice, sound effects and ambience come out of one generation, shaped together, with the character performance, the background texture and the timing all decided at once. ByteDance describes the target as end-to-end film-grade audio creation.

That framing is the whole point. Most audio tools hand you isolated outputs that you then assemble on a timeline. Seed Audio 1.0 is built to produce the finished moment.

What a sound scene actually means

Think about how a scene reaches a listener. There is a room before anyone speaks. There is pressure in a pause. A footstep lands just outside the frame, an alarm sits under the dialogue, and a cue arrives at exactly the right beat.

Producing that today usually means creating each layer separately, placing it on a timeline, adjusting timing and mix, then repeating until it feels right. Seed Audio 1.0 models those elements as parts of the same scene, so a voice learns how it sits in an environment, a sound cue learns how it supports a line, and timing shapes the emotion rather than being bolted on afterwards.

What a scene prompt looks like

The prompt is not a script with settings attached. It is a description of the moment, written the way you would brief a sound designer. ByteDance’s own example runs like this:

Try this prompt

Inside a huge football stadium, with the deafening roar of tens of thousands of fans throughout the background. The commentator (middle-aged male, British accent, rich and penetrating voice, classic sports commentary, extremely exhilarated) shouts in a rapid, soaring, full-throated tone: "OH, HE SCORES!!! WHAT A GOAL!" He draws out the word "GOAL" with a voice slightly hoarse from excitement, and the crowd's cheering erupts at the moment of the goal and continues to the end.

Read what is being specified there. The setting and its background noise. The speaker’s age, accent and vocal quality. The emotional state. The delivery on a single word. And the timing of the crowd, which erupts on the goal and holds to the end. All of that is one prompt, and all of it arrives as one piece of audio.

Three ways to prompt it

How you generate depends on what you hand over, and the model works out the mode from that.

Text only. Write the scene and it gets built from the description alone. The prompt runs to 3,000 characters, which is plenty of room to direct rather than just supply lines.

With reference audio. Attach up to three clips, each up to 30 seconds, and call them in the prompt as @Audio1, @Audio2 and @Audio3 in upload order. That is how you put a specific voice, or several, into a scene you are describing in words.

With a reference image. Hand it a single picture along with the text to be spoken. Note the trade-off: in this mode the prompt carries only the words, so you give up the scene direction you get in text mode.

Timing you can specify, down to 100 milliseconds

Seed Audio 1.0 accepts timing instructions in the prompt itself, accurate to 100 millisecond intervals, turning a creative idea into a structured timeline that says when each element should enter. That precision currently applies to character dialogue, so you can pin a line to an exact mark.

That sounds like a small detail until you are dubbing video. Matching a translated line to a mouth, dropping a re-voiced take into an existing edit, or hitting a beat in an advert are all jobs where timing is the deliverable, not a nice-to-have. Specifying it in the prompt removes the nudging-clips-on-a-timeline step entirely.

Building a voice from a description, a sample, or both

Generation is zero-shot, so no separate model gets trained for each speaker. A voice comes from a written description on its own, or from a description paired with a reference sample when you want finer control, which makes a new voice a prompt rather than a project. Reference clips accept wav, mp3, pcm and ogg_opus.

A character often has to move from calm narration to urgency, or from restrained reporting to full dramatic delivery, and still sound like the same person. Because the model learned voices inside scenes, alongside emotion and pacing and surrounding sound, it has more room to shift delivery without losing the identity underneath.

For longer pieces, it generates up to two minutes in a single pass and supports continuation from there, which is what keeps a character recognisable across an extended scene. The prompt itself runs to 3,000 characters.

One character, twenty languages

Seed Audio 1.0 covers 20 languages: English, Chinese, Japanese, Korean, Indonesian, German, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish and Swedish, plus Mexican and Castilian Spanish as separate options and Brazilian Portuguese.

The useful part is not the count, it is what travels and what adapts. The same voice transfers across languages, while its rhythm, stress, pauses and emotion shift to follow the expressive habits of the target language. A character stays recognisable without sounding like someone reading a foreign script phonetically. For game localisation, a campaign running in several markets, or a multilingual podcast, that means extending one voice rather than rebuilding the audio pipeline per language.

The controls you get on the output

Beyond the prompt, three sliders shape the finished audio: speech rate from half speed to double, volume across the same range, and pitch up or down twelve steps. Output comes as wav, mp3, pcm or ogg_opus.

Two extras are worth knowing. Turning on subtitles returns word-level timestamps, giving you the start and end of every word in milliseconds, which is what you need to caption or sync the audio without transcribing it again. And generated audio can carry a watermark, either an audible rhythm marker at the end of the file or hidden metadata in the header.

How well it holds up

ByteDance ran evaluations across three areas, and the numbers are worth knowing before you plan around the model.

Across a wide spread of scenarios, covering film and television, short drama, animation, podcast dialogue, live commerce, online content, stage plays and straight text to speech, more than 90 percent of generations came out usable. On multilingual output, naturalness scored above 4.0 for most languages. On following complex instructions, every language tested scored above 3.5 except Vietnamese.

Read that last one as a practical constraint. Vietnamese output is likely to need more attempts and closer checking than the rest.

What it is built for

The formats where sound has to carry the scene are the obvious fit: narrative audio, scripted dialogue, short-form video, advertising, game content and podcast-style production. The timing control makes video dubbing and re-voicing a particular strength, since those are jobs defined by hitting marks.

Straight single-voice narration is not where the model distinguishes itself. It will do it well, but nothing about a plain voiceover uses the scene modelling, the timing control or the cross-language identity work that make Seed Audio 1.0 different.

How it works underneath

Two problems have to be solved at once. At the language level, the model needs to understand who is speaking, what they feel, when each line and cue should land, and how the moment develops. At the acoustic level, it has to hold on to the details that make audio convincing: speaker identity, prosody, texture, impact, ambience and spatial continuity.

Rather than treating each audio type as a separate task, a single acoustic encoder captures voice, effects and ambience as parts of one coherent scene and maps them into a shared representation. A language model turns creative intent into scene-level controls on top of that, feeding a diffusion-based generator that renders the final audio in a high-fidelity latent space.

Where it goes next

Sound effects, ambience and music are already generated as part of the scene. What is not there yet is the same millisecond-level timing control over them, and extending that is what ByteDance has said comes next. Also on the roadmap: video as an input reference, long-form and multitrack generation, and controllable multilingual translation that manages expression, timing and duration across languages.

Seed Audio in the Picsart AI Playground

Seed Audio is available in the Picsart AI Playground as two separate models, and which one you want comes down to language.

Seed Audio synthesizes natural English or Chinese speech. Seed Audio Multilingual does the same across 20 languages. Both work the same way otherwise: pick a named voice, or clone one from a reference recording, and the model generates from there.

Voice cloning from reference audio is the shared headline feature, and it is what makes these two worth reaching for over a straightforward text-to-speech model. If a project needs one consistent character voice rather than a stock one, that is the reason to start here.

The Playground carries the pair alongside the rest of the audio catalogue, which spans text to speech, music, sound effects, voice changing, dubbing and audio cleanup from several providers. Everything runs in the same place with one credit balance, so moving between a voiceover, a backing track and a sound effect does not mean moving between tools.

Get answers to common questions

Seed Audio 1.0 is ByteDance’s audio creation model, released in July 2026. It generates voice, sound effects and ambience inside one unified framework, so a single prompt produces a coordinated scene rather than a set of separate clips.