From Speech to Audio Creation | Introducing the Seed Audio 1.0 Audio Creation Model

From Speech to Audio Creation | Introducing the Seed Audio 1.0 Audio Creation Model

Date

2026-07-20

Category

Models

Full-scene audio generation for creators

Creators rarely imagine sound as a set of isolated files.

They hear the room before a line is spoken. They hear the pressure in a pause, the footstep just outside the frame, the alarm buried under the dialogue, the background that makes a place feel real, and the cue that arrives at exactly the right moment.

Seed Audio 1.0 is built around that unit of creation: the scene.

It is an audio creation model that generates speech, sound effects, ambience, and other scene-level audio elements within one unified framework. Given a prompt, Seed Audio 1.0 can produce coordinated audio with character performance, timing, background texture, and sound design shaped together.

That scene-level approach also changes what the model can do with speech itself. Even when generating voice alone, Seed Audio 1.0 can produce a broader expressive range — from restrained narration to heightened emotion, from natural conversation to dramatic performance — while keeping the voice identity consistent.

The goal is to make audio generation feel less like assembling clips, and more like directing a sound scene.

Seed Audio 1.0 is now available through BytePlus.

Project page: https://seed.bytedance.com/en/seedaudio1_0

Try it now: https://console.byteplus.com/voice/new/setting/activate?projectName=default

Why scenes matter

Audio generation has made fast progress across individual capabilities: voice synthesis, sound effects, ambience, and multilingual speech. But creative work often depends on how these elements interact.

A voice line needs the right performance. A sound effect needs the right moment. Ambience needs to hold the space without overpowering the scene. Background texture needs to support emotion without breaking the pacing.

Today, much of that work still happens manually: creating each layer, placing it on a timeline, adjusting the timing and mix, and repeating until the scene feels right.

Seed Audio 1.0 explores a more integrated path. It models multiple audio elements as parts of the same sound scene, helping creators move from isolated outputs to complete sound moments — scenes that carry dialogue, atmosphere, action, and emotion in one coherent experience.

What Seed Audio 1.0 can do

Generate coordinated sound scenes and control dialogue timing

A prompt can describe the speaker, the emotional tone, the line delivery, the surrounding environment, key sound effects, and how the scene should unfold.

Seed Audio 1.0 models them under a shared scene context, allowing dialogue, ambience, and sound cues to align more naturally in timing, tone, and acoustic space.

Seed Audio 1.0 can create complex audio scenes from a single prompt.

This makes Seed Audio 1.0 useful for narrative audio, scripted dialogue, short-form video, advertising, game content, podcast-style production, and other formats where sound needs to carry the scene.

Seed Audio 1.0 supports prompt-level timing control for character dialogue, with current timing precision at 100 ms intervals. Creators can specify when lines should enter, making the model useful for video dubbing, re-voicing, advertising, and scripted audio where timing is part of the deliverable.

Generate voices with range and consistency

Seed Audio 1.0 lets creators shape a voice from a text description, an authorized reference sample, or both — without training a separate model for each speaker.

But a voice is more than a timbre. In real creative work, the same character may need to move from calm narration to urgency, from restrained reporting to dramatic dialogue, while still sounding like the same person.

Because Seed Audio 1.0 is trained not only as a speech generator, but as an audio creation model, it learns voices in the context of scenes: emotion, pacing, surrounding sound, and narrative intent. Even in speech-only generation, this gives the model more room to shape delivery across emotion, prosody, rhythm, and speaking style.

For longer content, Seed Audio 1.0 can generate up to two minutes of audio in a single pass and supports further continuation, helping a character stay recognizable across longer scenes and repeated extensions.

Create multilingual audio with natural expression

For global creators, multilingual audio is not a translation problem alone. A character voice has to carry the same identity across languages while adapting to the rhythm, pronunciation, pacing, and emotional habits of each local audience.

Seed Audio 1.0 supports audio generation across 20+ languages, including Chinese, English, Japanese, Korean, Spanish, Indonesian, German, French, Thai, and Vietnamese. The model can adapt a character voice across languages while preserving its recognizable timbre and performance style.

This is especially useful for global content rollout, game localization, brand campaigns, multilingual podcasts, and short-form video. A team can start from one creative concept or character voice and extend it across markets more efficiently, without rebuilding the entire audio production pipeline for every language.

A unified model for sound scenes

Full-scene audio generation is difficult because it requires two kinds of understanding at once.

At the language level, the model has to understand the scene: who is speaking, what emotion they carry, when each line or sound cue should enter, and how the moment should unfold. At the acoustic level, it has to preserve the details that make audio convincing: speaker identity, prosody, texture, impact, ambience, and spatial continuity.

Seed Audio 1.0 connects these two levels through a unified generation framework. A language model helps structure the creative intent into scene-level controls, while a diffusion-based acoustic generator produces the final audio in a high-fidelity latent space. This latent representation is designed to preserve acoustic detail while remaining controllable by semantic instructions such as character, emotion, timing, and scene context.

This is what allows Seed Audio 1.0 to model speech, sound effects, ambience, and other audio elements as parts of the same scene rather than as unrelated tasks. The model can learn how a voice sits in an environment, how a sound cue supports a line, and how timing changes the emotional shape of a scene.

Evaluation

We evaluated Seed Audio 1.0 across three areas: text-prompted voice generation, multi-scenario audio creation, and multilingual generation.

In A/B subjective evaluations, Seed Audio 1.0 showed clear preference gains in generating voices from text descriptions. It also improved both usability and high-quality output rates, suggesting stronger control over voice design and performance.

Text-to-Timbre Subjective Benchmark: Seed Audio 1.0 vs. Competitors

For multi-scenario audio creation, we tested Seed Audio 1.0 across a wide range of use cases, including film and television, TV programming, short drama, animation, podcast dialogue, live commerce, online content creation, stage performance, and text-to-speech. Across most scenarios, the usable audio rate exceeded 90%.

Conditional Audio-to-Audio Evaluation of Seed Audio 1.0 by Scenario

Conditional Audio-to-Audio Evaluation of Seed Audio 1.0 by Scenario

For multilingual generation, we used human evaluation to measure instruction following and audio naturalness. On audio naturalness, most languages achieved MOS scores above 4.0. On instruction following for complex audio generation tasks, all evaluated languages except Vietnamese scored above 3.5.

Seed Audio 1.0 Performance on T2A (Text-to-Audio) and A2A (Audio-to-Audio) Tasks

Seed Audio 1.0 Performance on T2A (Text-to-Audio) Tasks

Together, these results suggest that Seed Audio 1.0 can support a broad range of practical creative tasks, from expressive voice generation to complex multilingual audio scenes.

What comes next

Seed Audio 1.0 is an early step toward a broader form of audio creation.

Today, timing control is focused mainly on character dialogue. We will continue extending fine-grained control to more sound elements, including sound effects, ambience, and music. We will also keep improving voice consistency in longer and more complex scenes.

Looking ahead, we plan to support more input modalities, including video references, and to explore more advanced capabilities for long-form and multitrack audio generation. We will also explore controllable multilingual translation, so creators can better manage expression, timing, and duration across languages.

The long-term direction is clear: audio generation should become more expressive, more controllable, and more aligned with how creators actually think.

Seed Audio 1.0 is a step in that direction — helping creators turn imagined sound scenes into finished audio.