{"id":260860,"date":"2026-07-31T15:33:45","date_gmt":"2026-07-31T22:33:45","guid":{"rendered":"https:\/\/picsart.com\/blog\/?p=260860"},"modified":"2026-07-31T17:42:26","modified_gmt":"2026-08-01T00:42:26","slug":"seed-audio-1-0-explained","status":"publish","type":"post","link":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/","title":{"rendered":"Seed Audio 1.0 explained: one prompt for voice, effects and ambience"},"content":{"rendered":"<p>Seed Audio 1.0 is ByteDance&#8217;s audio creation model, released in July 2026, and it treats a piece of audio as a scene rather than a stack of separate files. Voice, sound effects and ambience come out of one generation, shaped together, with the character performance, the background texture and the timing all decided at once. ByteDance describes the target as end-to-end film-grade audio creation.<\/p>\n<p>That framing is the whole point. Most audio tools hand you isolated outputs that you then assemble on a timeline. Seed Audio 1.0 is built to produce the finished moment.<\/p>\n<h2><span id=\"What_a_sound_scene_actually_means\">What a sound scene actually means<\/span><\/h2>\n<p>Think about how a scene reaches a listener. There is a room before anyone speaks. There is pressure in a pause. A footstep lands just outside the frame, an alarm sits under the dialogue, and a cue arrives at exactly the right beat.<\/p>\n<p>Producing that today usually means creating each layer separately, placing it on a timeline, adjusting timing and mix, then repeating until it feels right. Seed Audio 1.0 models those elements as parts of the same scene, so a voice learns how it sits in an environment, a sound cue learns how it supports a line, and timing shapes the emotion rather than being bolted on afterwards.<\/p>\n<h2><span id=\"What_a_scene_prompt_looks_like\">What a scene prompt looks like<\/span><\/h2>\n<p>The prompt is not a script with settings attached. It is a description of the moment, written the way you would brief a sound designer. ByteDance&#8217;s own example runs like this:<\/p>\n<p><code><section class=\"try_prompt_block\" data-pulse-section-group=\"blog article\">\n    <h3 class=\"try_prompt_title\">Try this prompt<\/h3>\n\n    <div class=\"try_prompt_card\" data-pulse-section=\"blog article_try prompt\">\n        <p class=\"try_prompt_text\" id=\"try-prompt-6a6d7cbfc08f1-text\">Inside a huge football stadium, with the deafening roar of tens of thousands of fans throughout the background. The commentator (middle-aged male, British accent, rich and penetrating voice, classic sports commentary, extremely exhilarated) shouts in a rapid, soaring, full-throated tone: &quot;OH, HE SCORES!!! WHAT A GOAL!&quot; He draws out the word &quot;GOAL&quot; with a voice slightly hoarse from excitement, and the crowd&#039;s cheering erupts at the moment of the goal and continues to the end.<\/p>\n\n        <button type=\"button\"\n                class=\"try_prompt_copy\"\n                data-copy-target=\"#try-prompt-6a6d7cbfc08f1-text\"\n                aria-label=\"Copy prompt\"\n                title=\"Copy prompt\"\n                data-pulse-name=\"try prompt - copy\">\n            <img class=\"try_prompt_copy_icon\"\n                 src=\"https:\/\/cdn-cms-uploads.picsart.com\/cms-uploads\/79a9b2ac-f4c9-436b-b388-a05a484adf00.png\"\n                 alt=\"\"\n                 width=\"20\" height=\"20\"\n                 loading=\"lazy\" decoding=\"async\">\n            <svg class=\"try_prompt_check_icon\" viewBox=\"0 0 24 24\" width=\"20\" height=\"20\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" aria-hidden=\"true\">\n                <polyline points=\"5,12 10,17 19,7\"><\/polyline>\n            <\/svg>\n            <span class=\"try_prompt_copy_tooltip\" role=\"status\" aria-live=\"polite\">Copied<\/span>\n        <\/button>\n    <\/div>\n\n    <\/section>\n<\/code><\/p>\n<p><code><\/code>Read what is being specified there. The setting and its background noise. The speaker&#8217;s age, accent and vocal quality. The emotional state. The delivery on a single word. And the timing of the crowd, which erupts on the goal and holds to the end. All of that is one prompt, and all of it arrives as one piece of audio.<\/p>\n<h2><span id=\"Three_ways_to_prompt_it\">Three ways to prompt it<\/span><\/h2>\n<p>How you generate depends on what you hand over, and the model works out the mode from that.<\/p>\n<p><strong>Text only.<\/strong> Write the scene and it gets built from the description alone. The prompt runs to 3,000 characters, which is plenty of room to direct rather than just supply lines.<\/p>\n<p><strong>With reference audio.<\/strong> Attach up to three clips, each up to 30 seconds, and call them in the prompt as @Audio1, @Audio2 and @Audio3 in upload order. That is how you put a specific voice, or several, into a scene you are describing in words.<\/p>\n<p><strong>With a reference image.<\/strong> Hand it a single picture along with the text to be spoken. Note the trade-off: in this mode the prompt carries only the words, so you give up the scene direction you get in text mode.<\/p>\n<h2><span id=\"Timing_you_can_specify_down_to_100_milliseconds\">Timing you can specify, down to 100 milliseconds<\/span><\/h2>\n<p>Seed Audio 1.0 accepts timing instructions in the prompt itself, accurate to 100 millisecond intervals, turning a creative idea into a structured timeline that says when each element should enter. That precision currently applies to character dialogue, so you can pin a line to an exact mark.<\/p>\n<p>That sounds like a small detail until you are dubbing video. Matching a translated line to a mouth, dropping a re-voiced take into an existing edit, or hitting a beat in an advert are all jobs where timing is the deliverable, not a nice-to-have. Specifying it in the prompt removes the nudging-clips-on-a-timeline step entirely.<\/p>\n<h2><span id=\"Building_a_voice_from_a_description_a_sample_or_both\">Building a voice from a description, a sample, or both<\/span><\/h2>\n<p>Generation is zero-shot, so no separate model gets trained for each speaker. A voice comes from a written description on its own, or from a description paired with a reference sample when you want finer control, which makes a new voice a prompt rather than a project. Reference clips accept wav, mp3, pcm and ogg_opus.<\/p>\n<p>A character often has to move from calm narration to urgency, or from restrained reporting to full dramatic delivery, and still sound like the same person. Because the model learned voices inside scenes, alongside emotion and pacing and surrounding sound, it has more room to shift delivery without losing the identity underneath.<\/p>\n<p>For longer pieces, it generates up to two minutes in a single pass and supports continuation from there, which is what keeps a character recognisable across an extended scene. The prompt itself runs to 3,000 characters.<\/p>\n<h2><span id=\"One_character_twenty_languages\">One character, twenty languages<\/span><\/h2>\n<p>Seed Audio 1.0 covers 20 languages: English, Chinese, Japanese, Korean, Indonesian, German, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish and Swedish, plus Mexican and Castilian Spanish as separate options and Brazilian Portuguese.<\/p>\n<p>The useful part is not the count, it is what travels and what adapts. The same voice transfers across languages, while its rhythm, stress, pauses and emotion shift to follow the expressive habits of the target language. A character stays recognisable without sounding like someone reading a foreign script phonetically. For game localisation, a campaign running in several markets, or a multilingual podcast, that means extending one voice rather than rebuilding the audio pipeline per language.<\/p>\n<h2><span id=\"The_controls_you_get_on_the_output\">The controls you get on the output<\/span><\/h2>\n<p>Beyond the prompt, three sliders shape the finished audio: speech rate from half speed to double, volume across the same range, and pitch up or down twelve steps. Output comes as wav, mp3, pcm or ogg_opus.<\/p>\n<p>Two extras are worth knowing. Turning on subtitles returns word-level timestamps, giving you the start and end of every word in milliseconds, which is what you need to caption or sync the audio without transcribing it again. And generated audio can carry a watermark, either an audible rhythm marker at the end of the file or hidden metadata in the header.<\/p>\n<h2><span id=\"How_well_it_holds_up\">How well it holds up<\/span><\/h2>\n<p>ByteDance ran evaluations across three areas, and the numbers are worth knowing before you plan around the model.<\/p>\n<p>Across a wide spread of scenarios, covering film and television, short drama, animation, podcast dialogue, live commerce, online content, stage plays and straight text to speech, more than 90 percent of generations came out usable. On multilingual output, naturalness scored above 4.0 for most languages. On following complex instructions, every language tested scored above 3.5 except Vietnamese.<\/p>\n<p>Read that last one as a practical constraint. Vietnamese output is likely to need more attempts and closer checking than the rest.<\/p>\n<h2><span id=\"What_it_is_built_for\">What it is built for<\/span><\/h2>\n<p>The formats where sound has to carry the scene are the obvious fit: narrative audio, scripted dialogue, short-form video, advertising, game content and podcast-style production. The timing control makes video dubbing and re-voicing a particular strength, since those are jobs defined by hitting marks.<\/p>\n<p>Straight single-voice narration is not where the model distinguishes itself. It will do it well, but nothing about a plain voiceover uses the scene modelling, the timing control or the cross-language identity work that make Seed Audio 1.0 different.<\/p>\n<h2><span id=\"How_it_works_underneath\">How it works underneath<\/span><\/h2>\n<p>Two problems have to be solved at once. At the language level, the model needs to understand who is speaking, what they feel, when each line and cue should land, and how the moment develops. At the acoustic level, it has to hold on to the details that make audio convincing: speaker identity, prosody, texture, impact, ambience and spatial continuity.<\/p>\n<p>Rather than treating each audio type as a separate task, a single acoustic encoder captures voice, effects and ambience as parts of one coherent scene and maps them into a shared representation. A language model turns creative intent into scene-level controls on top of that, feeding a diffusion-based generator that renders the final audio in a high-fidelity latent space.<\/p>\n<h2><span id=\"Where_it_goes_next\">Where it goes next<\/span><\/h2>\n<p>Sound effects, ambience and music are already generated as part of the scene. What is not there yet is the same millisecond-level timing control over them, and extending that is what ByteDance has said comes next. Also on the roadmap: video as an input reference, long-form and multitrack generation, and controllable multilingual translation that manages expression, timing and duration across languages.<\/p>\n<h2><span id=\"Seed_Audio_in_the_Picsart_AI_Playground\">Seed Audio in the Picsart AI Playground<\/span><\/h2>\n<p>Seed Audio is available in the <a href=\"https:\/\/picsart.com\/ai-playground\/\">Picsart AI Playground<\/a> as two separate models, and which one you want comes down to language.<\/p>\n<p><strong>Seed Audio<\/strong> synthesizes natural English or Chinese speech. <strong>Seed Audio Multilingual<\/strong> does the same across 20 languages. Both work the same way otherwise: pick a named voice, or clone one from a reference recording, and the model generates from there.<\/p>\n<p>Voice cloning from reference audio is the shared headline feature, and it is what makes these two worth reaching for over a straightforward text-to-speech model. If a project needs one consistent character voice rather than a stock one, that is the reason to start here.<\/p>\n<p>The Playground carries the pair alongside the rest of the audio catalogue, which spans text to speech, music, sound effects, voice changing, dubbing and audio cleanup from several providers. Everything runs in the same place with one credit balance, so moving between a voiceover, a backing track and a sound effect does not mean moving between tools.<\/p>\n<section class=\"section_faq\" id=\"faq-faq-6a6d7cbfc0bf9\">\n            <h2 class=\"faq_title\" id=\"Get_answers_to_common_questions\">Get answers to common questions<\/h2>\n    \n    <div class=\"faq_items\">\n                    <div class=\"faq_item faq_item--active\">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"true\">\n                    <span class=\"faq_question_text\">What is Seed Audio 1.0?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"false\">\n                    <div class=\"faq_answer_content\"><p>Seed Audio 1.0 is ByteDance&#8217;s audio creation model, released in July 2026. It generates voice, sound effects and ambience inside one unified framework, so a single prompt produces a coordinated scene rather than a set of separate clips.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">How many languages does Seed Audio 1.0 support?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>20, including English, Chinese, Japanese, Korean, Indonesian, German, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish and Swedish. Mexican and Castilian Spanish are listed separately, and Portuguese is Brazilian.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">Can Seed Audio 1.0 clone a voice?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>Yes, and generation is zero-shot, so no separate model is trained per speaker. A voice can come from a written description alone, or from a description paired with up to three reference clips of 30 seconds each for finer control.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">Can Seed Audio 1.0 generate audio from an image?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>Yes. Supplying a single reference image along with the text to be spoken generates audio from that image. Image references cannot be combined with audio references, and in that mode the prompt carries only the words to be said.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">How long can a single generation be?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>Up to two minutes in one pass, with continuation supported beyond that, which helps a character stay consistent across a longer scene. The prompt itself runs to 3,000 characters.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">What is the difference between Seed Audio and Seed Audio Multilingual?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>Seed Audio covers English and Chinese. Seed Audio Multilingual covers 20 languages. Both let you pick a named voice or clone one from a reference recording.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n            <\/div>\n<\/section>\n\n<script type=\"application\/ld+json\">\n{\n    \"@context\": \"https:\/\/schema.org\",\n    \"@type\": \"FAQPage\",\n    \"mainEntity\": [\n        {\n            \"@type\": \"Question\",\n            \"name\": \"What is Seed Audio 1.0?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Seed Audio 1.0 is ByteDance&#8217;s audio creation model, released in July 2026. It generates voice, sound effects and ambience inside one unified framework, so a single prompt produces a coordinated scene rather than a set of separate clips.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"How many languages does Seed Audio 1.0 support?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"20, including English, Chinese, Japanese, Korean, Indonesian, German, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish and Swedish. Mexican and Castilian Spanish are listed separately, and Portuguese is Brazilian.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Can Seed Audio 1.0 clone a voice?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Yes, and generation is zero-shot, so no separate model is trained per speaker. A voice can come from a written description alone, or from a description paired with up to three reference clips of 30 seconds each for finer control.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Can Seed Audio 1.0 generate audio from an image?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Yes. Supplying a single reference image along with the text to be spoken generates audio from that image. Image references cannot be combined with audio references, and in that mode the prompt carries only the words to be said.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"How long can a single generation be?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Up to two minutes in one pass, with continuation supported beyond that, which helps a character stay consistent across a longer scene. The prompt itself runs to 3,000 characters.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"What is the difference between Seed Audio and Seed Audio Multilingual?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Seed Audio covers English and Chinese. Seed Audio Multilingual covers 20 languages. Both let you pick a named voice or clone one from a reference recording.\"\n            }\n        }\n    ]\n}<\/script>\n\n<script>\n(function() {\n    var container = document.getElementById('faq-faq-6a6d7cbfc0bf9');\n    if (!container) return;\n\n    var items = container.querySelectorAll('.faq_item');\n    items.forEach(function(item) {\n        var button = item.querySelector('.faq_question');\n        var answer = item.querySelector('.faq_answer');\n        if (!button || !answer) return;\n\n        button.addEventListener('click', function() {\n            var isActive = item.classList.contains('faq_item--active');\n\n            if (isActive) {\n                item.classList.remove('faq_item--active');\n                button.setAttribute('aria-expanded', 'false');\n                answer.setAttribute('aria-hidden', 'true');\n                answer.setAttribute('data-collapsed', '');\n            } else {\n                items.forEach(function(other) {\n                    var otherBtn = other.querySelector('.faq_question');\n                    var otherAnswer = other.querySelector('.faq_answer');\n                    other.classList.remove('faq_item--active');\n                    if (otherBtn) otherBtn.setAttribute('aria-expanded', 'false');\n                    if (otherAnswer) {\n                        otherAnswer.setAttribute('aria-hidden', 'true');\n                        otherAnswer.setAttribute('data-collapsed', '');\n                    }\n                });\n                item.classList.add('faq_item--active');\n                button.setAttribute('aria-expanded', 'true');\n                answer.removeAttribute('data-collapsed');\n                answer.setAttribute('aria-hidden', 'false');\n            }\n        });\n    });\n})();\n<\/script>\n\n","protected":false},"excerpt":{"rendered":"<p>Seed Audio 1.0 is ByteDance&#8217;s audio creation model, released in July 2026, and it treats a piece of audio as a scene rather than a stack of separate files. Voice, sound effects and ambience come out of one generation, shaped together, with the character performance, the background texture and the timing all decided at once. &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Seed Audio 1.0 explained: one prompt for voice, effects and ambience&#8221;<\/span><\/a><\/p>\n","protected":false},"author":146,"featured_media":237926,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_yoast_wpseo_title":"What is Seed Audio 1.0: ByteDance's audio creation model","_yoast_wpseo_metadesc":"Seed Audio 1.0 explained: ByteDance's model makes voice, sound effects and ambience as one scene, in 20 languages, with voice cloning from a recording.","faq_show":true,"faq_enable_schema":true,"how_to_show":false,"how_to_show_on_single":false,"how_to_enable_schema":false,"how_to_is_upload":false,"faq_title":"Get answers to common questions","how_to_title":"","how_to_layout":"","how_to_cta_text":"","how_to_cta_url":"","how_to_image_alt":"","how_to_display_image":0,"faq_items":null,"how_to_steps":[],"prompt_box_show":false,"prompt_box_placeholder":"","prompt_box_deeplink":"","prompt_box_submit_label":"","footnotes":""},"categories":[3181,1669],"tags":[],"class_list":["post-260860","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai","category-inspiration","entry"],"acf":{"footer_banner_name":"Start your design in Picsart","footer_banner_link_":"\/","footer_banner_button_text_":"Get Started","faq_show":true,"faq_title":"Get answers to common questions","faq_enable_schema":true,"faq_items":[{"question":"What is Seed Audio 1.0?","answer":"Seed Audio 1.0 is ByteDance's audio creation model, released in July 2026. It generates voice, sound effects and ambience inside one unified framework, so a single prompt produces a coordinated scene rather than a set of separate clips."},{"question":"How many languages does Seed Audio 1.0 support?","answer":"20, including English, Chinese, Japanese, Korean, Indonesian, German, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish and Swedish. Mexican and Castilian Spanish are listed separately, and Portuguese is Brazilian."},{"question":"Can Seed Audio 1.0 clone a voice?","answer":"Yes, and generation is zero-shot, so no separate model is trained per speaker. A voice can come from a written description alone, or from a description paired with up to three reference clips of 30 seconds each for finer control."},{"question":"Can Seed Audio 1.0 generate audio from an image?","answer":"Yes. Supplying a single reference image along with the text to be spoken generates audio from that image. Image references cannot be combined with audio references, and in that mode the prompt carries only the words to be said."},{"question":"How long can a single generation be?","answer":"Up to two minutes in one pass, with continuation supported beyond that, which helps a character stay consistent across a longer scene. The prompt itself runs to 3,000 characters."},{"question":"What is the difference between Seed Audio and Seed Audio Multilingual?","answer":"Seed Audio covers English and Chinese. Seed Audio Multilingual covers 20 languages. Both let you pick a named voice or clone one from a reference recording."}],"how_to_show":false,"how_to_show_on_single":false,"how_to_title":"","how_to_layout":"default","how_to_steps":null,"how_to_enable_schema":false,"how_to_is_upload":true,"how_to_cta_text":"","how_to_cta_url":"https:\/\/picsart.com\/create\/editor","how_to_display_image":null,"how_to_image_alt":"","prompt_box_show":false,"prompt_box_placeholder":"","prompt_box_deeplink":"https:\/\/picsart.com\/create\/editor?category=miniapps&app=com.picsart.aura","prompt_box_submit_label":"Create","try_prompt_show":true,"try_prompt_title":"Try this prompt","try_prompt_text":"Inside a huge football stadium, with the deafening roar of tens of thousands of fans throughout the background. The commentator (middle-aged male, British accent, rich and penetrating voice, classic sports commentary, extremely exhilarated) shouts in a rapid, soaring, full-throated tone: \"OH, HE SCORES!!! WHAT A GOAL!\" He draws out the word \"GOAL\" with a voice slightly hoarse from excitement, and the crowd's cheering erupts at the moment of the goal and continues to the end.","try_prompt_deeplink":"","tips_show":false,"tips_title":"Tips for best results","tips_items":null,"cta_banner_show":false,"cta_banner_title":"Need more space?","cta_banner_subtitle":"Extend any image in any direction with AI.","cta_banner_button_label":"Expand image","cta_banner_button_url":"","related_tools_title":"Related tools","related_tools_items":null,"post_level":""},"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v25.5 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>What is Seed Audio 1.0: ByteDance&#039;s audio creation model<\/title>\n<meta name=\"description\" content=\"Seed Audio 1.0 explained: ByteDance&#039;s model makes voice, sound effects and ambience as one scene, in 20 languages, with voice cloning from a recording.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"What is Seed Audio 1.0: ByteDance&#039;s audio creation model\" \/>\n<meta property=\"og:description\" content=\"Seed Audio 1.0 explained: ByteDance&#039;s model makes voice, sound effects and ambience as one scene, in 20 languages, with voice cloning from a recording.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/\" \/>\n<meta property=\"og:site_name\" content=\"Picsart Blog\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/picsart\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-31T22:33:45+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-01T00:42:26+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/cdnblog.picsart.com\/2025\/09\/CR5955-_collaboration-with-freepik-and-bytedance_1200x800_07.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"2400\" \/>\n\t<meta property=\"og:image:height\" content=\"1600\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Julia Tovmasyan\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@PicsArtStudio\" \/>\n<meta name=\"twitter:site\" content=\"@PicsArtStudio\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Julia Tovmasyan\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"7 minutes\" \/>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"What is Seed Audio 1.0: ByteDance's audio creation model","description":"Seed Audio 1.0 explained: ByteDance's model makes voice, sound effects and ambience as one scene, in 20 languages, with voice cloning from a recording.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/","og_locale":"en_US","og_type":"article","og_title":"What is Seed Audio 1.0: ByteDance's audio creation model","og_description":"Seed Audio 1.0 explained: ByteDance's model makes voice, sound effects and ambience as one scene, in 20 languages, with voice cloning from a recording.","og_url":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/","og_site_name":"Picsart Blog","article_publisher":"https:\/\/www.facebook.com\/picsart","article_published_time":"2026-07-31T22:33:45+00:00","article_modified_time":"2026-08-01T00:42:26+00:00","og_image":[{"width":2400,"height":1600,"url":"https:\/\/cdnblog.picsart.com\/2025\/09\/CR5955-_collaboration-with-freepik-and-bytedance_1200x800_07.jpg","type":"image\/jpeg"}],"author":"Julia Tovmasyan","twitter_card":"summary_large_image","twitter_creator":"@PicsArtStudio","twitter_site":"@PicsArtStudio","twitter_misc":{"Written by":"Julia Tovmasyan","Est. reading time":"7 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/#article","isPartOf":{"@id":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/"},"author":{"name":"Julia Tovmasyan","@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/person\/74b70f3125250c23596a5306775b702d"},"headline":"Seed Audio 1.0 explained: one prompt for voice, effects and ambience","datePublished":"2026-07-31T22:33:45+00:00","dateModified":"2026-08-01T00:42:26+00:00","mainEntityOfPage":{"@id":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/"},"wordCount":1475,"publisher":{"@id":"https:\/\/picsart.com\/blog\/ko\/#organization"},"image":{"@id":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/#primaryimage"},"thumbnailUrl":"https:\/\/cdnblog.picsart.com\/2025\/09\/CR5955-_collaboration-with-freepik-and-bytedance_1200x800_07.jpg","articleSection":["AI","Inspirational"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/","url":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/","name":"What is Seed Audio 1.0: ByteDance's audio creation model","isPartOf":{"@id":"https:\/\/picsart.com\/blog\/ko\/#website"},"primaryImageOfPage":{"@id":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/#primaryimage"},"image":{"@id":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/#primaryimage"},"thumbnailUrl":"https:\/\/cdnblog.picsart.com\/2025\/09\/CR5955-_collaboration-with-freepik-and-bytedance_1200x800_07.jpg","datePublished":"2026-07-31T22:33:45+00:00","dateModified":"2026-08-01T00:42:26+00:00","description":"Seed Audio 1.0 explained: ByteDance's model makes voice, sound effects and ambience as one scene, in 20 languages, with voice cloning from a recording.","breadcrumb":{"@id":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/#primaryimage","url":"https:\/\/cdnblog.picsart.com\/2025\/09\/CR5955-_collaboration-with-freepik-and-bytedance_1200x800_07.jpg","contentUrl":"https:\/\/cdnblog.picsart.com\/2025\/09\/CR5955-_collaboration-with-freepik-and-bytedance_1200x800_07.jpg","width":2400,"height":1600,"caption":"Picsart collaboration with ByteDance to release seedream 4"},{"@type":"BreadcrumbList","@id":"https:\/\/picsart.com\/blog\/seed-audio-1-0-explained\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/picsart.com\/blog\/"},{"@type":"ListItem","position":2,"name":"Seed Audio 1.0 explained: one prompt for voice, effects and ambience"}]},{"@type":"WebSite","@id":"https:\/\/picsart.com\/blog\/ko\/#website","url":"https:\/\/picsart.com\/blog\/ko\/","name":"Picsart Blog","description":"Keep up with the latest news in photo editing, digital photography, and art trends.","publisher":{"@id":"https:\/\/picsart.com\/blog\/ko\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/picsart.com\/blog\/ko\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/picsart.com\/blog\/ko\/#organization","name":"PicsArt Inc.","url":"https:\/\/picsart.com\/blog\/ko\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/logo\/image\/","url":"https:\/\/cdnblog.picsart.com\/2016\/02\/PicsArt-logo.png","contentUrl":"https:\/\/cdnblog.picsart.com\/2016\/02\/PicsArt-logo.png","width":195,"height":43,"caption":"PicsArt Inc."},"image":{"@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/picsart","https:\/\/x.com\/PicsArtStudio","https:\/\/www.instagram.com\/picsart","https:\/\/www.linkedin.com\/company\/picsart-photo-studio","https:\/\/www.pinterest.com\/picsart"]},{"@type":"Person","@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/person\/74b70f3125250c23596a5306775b702d","name":"Julia Tovmasyan","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/person\/image\/","url":"https:\/\/cdnblog.picsart.com\/2026\/03\/3285C16C-FD87-4868-A2F0-04B6A0815CE1-150x150.jpg","contentUrl":"https:\/\/cdnblog.picsart.com\/2026\/03\/3285C16C-FD87-4868-A2F0-04B6A0815CE1-150x150.jpg","caption":"Julia Tovmasyan"}}]}},"featured_image":{"url":"https:\/\/cdnblog.picsart.com\/2025\/09\/CR5955-_collaboration-with-freepik-and-bytedance_1200x800_07.jpg","dimensions":[]},"_links":{"self":[{"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/posts\/260860","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/users\/146"}],"replies":[{"embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/comments?post=260860"}],"version-history":[{"count":7,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/posts\/260860\/revisions"}],"predecessor-version":[{"id":260867,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/posts\/260860\/revisions\/260867"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/media\/237926"}],"wp:attachment":[{"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/media?parent=260860"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/categories?post=260860"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/tags?post=260860"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}