Seed Audio 2.0: Six-Minute Scenes, Video Dubbing, and Separate Stems
What the published Seed Audio 2.0 spec adds to Seed Audio 1.0: longer scenes, video input, six voice references, and stems you can still edit.

Seed Audio 1.0 already treats a soundtrack as one scene: dialogue, room tone, effects, and music from a single brief. The follow-up now written up in public, Seed Audio 2.0, is aimed at the jobs that still break that pass — a cut longer than two minutes, a picture the sound has to follow, and a mix an editor needs to pull apart.
This note was checked on September 28, 2026. It summarizes ByteDance’s Seed Audio 1.0 announcement and the capability list published for Seed Audio 2.0. A control counts as available only when the studio you are using actually exposes it.
What Seed Audio 1.0 already does
On July 20, 2026, ByteDance’s Seed team published From Speech to Audio Creation. The model is built around a scene, not a lone voice file. A prompt can name the speaker, the emotion, the line, the room, and the cues. Dialogue timing is described in 100-millisecond steps. One pass runs about two minutes, with continuation when a scene has to go on. The post lists 20-plus languages, including Chinese, English, Japanese, Korean, Spanish, Indonesian, German, French, Thai, and Vietnamese.
The same post is plain about what comes next: finer control over effects, ambience, and music; video as an input; longer, multitrack audio; and more controllable multilingual performance. Those are the gaps the Seed Audio 2.0 spec is written to close.
What the Seed Audio 2.0 spec adds
Seed Audio 2.0 describes four ways to start a job. The names are the inputs, not four different products.
- T2A — text to audio. A script or scene brief, with no recording attached.
- TA2A — text plus reference audio. Use this when a voice has to stay recognizable.
- TV2A — text plus video. The picture is the timeline the dub and the bed are supposed to follow.
- TAV2A — text, reference audio, and video together. A known cast plus a locked cut.
Set beside the 1.0 model ByteDance has actually published, the 2.0 product page describes this delta:
| Workflow | Seed Audio 1.0 | Seed Audio 2.0 spec |
|---|---|---|
| Length of one pass | About 2 minutes, plus continuation | Up to 6 minutes |
| Reference clips | About 3 | Up to 6 |
| Inputs | Text, reference audio, and an optional still image | Text, reference audio, and reference video: T2A, TA2A, TV2A, TAV2A |
| Languages | 20-plus in the official 1.0 post | 30 on the 2.0 product page |
| Editability | Dialogue timing, then one mixed file | Independent stems for dialogue, music, ambience, and effects, plus timestamps |
| Best fit | Short scenes and sound sketches | Longer drama, ads, games, dubs, and crowded casts |
Secondary write-ups repeat the same four changes: six minutes, six references, separate stems, and cues you can place in time. They are useful as a cross-check, not as a second source of truth. Version names in the wild are less stable. Some directories have filed a near-identical ByteDance follow-up as Seed Audio 1.5, and at least one page that used the 2.0 name later called that label a mistake. Until ByteDance posts a versioned release note, read the list on Seed Audio 2.0 as the working spec, and assume a live generator is still 1.0 unless that page says the new model is on.
Why stems matter more than a longer file
A six-minute mix helps only if you can still edit it. A short-drama editor who needs the rain lower under a whisper, or a localization producer who must replace one market’s voice without rebuilding the city bed, cannot do that from a single bounce. Separate dialogue, music, ambience, and effects tracks are the difference between a temp soundtrack and a session.
Timestamps are the other half. Seed Audio 1.0 already places dialogue on a rough grid. The 2.0 brief extends that idea to effects and music, so a door, a sting, or a line pickup can be requested at a marked time instead of hoped for somewhere in the take.
Six references matter for the same reason a three-voice scene usually collapses. A lead, a foil, and a bit of crowd texture already use the 1.0 budget. The extra slots are for keeping a larger cast distinct, not for uploading six random songs and hoping the model averages them.
Where that spec is aimed
Public product copy points Seed Audio 2.0 at work that 1.0 can sketch but not finish in one job:
- Short drama. Several speakers in one take, with footsteps and a sting left on their own tracks.
- Ads. Voice, product Foley, and a music lift generated together, then split for a legal recut or another market.
- Games. Barks, loops, and hits briefed as one scene, then handed to a mixer as stems.
- Dubbing. Video context plus the wider language set, for teams that currently re-record every market by hand.
A single narrator, a one-line voiceover, or an eight-second social hook may still be simpler on Seed Audio 1.0. A longer ceiling does not fix a vague brief.
A brief you can rehearse before launch
The studio on Seed Audio 2.0 still generates with Seed Audio 1.0 while the new model is marked coming soon. That is a useful limit. Write the 2.0 brief anyway, then listen to the part 1.0 can already make: speakers, room, and cue order.
Mode: TAV2A
Length: 90 seconds
Language: English
Speakers:
- Captain, reference A: tired, never shouts
- Engineer, reference B: quicker, technical
Picture: locked bridge-alarm cut
Stems: dialogue, music, ambience, effects
Cues:
- 00:04 door seal
- 00:12 captain line as the alarm climbs
- 00:28 engines drop and the strings thin
Listen for three failures: voices that trade identity, music that covers a word, and effects that land late. Those are the notes stems and timestamps are supposed to reduce. You can draft the same scene shape in the Seed Audio 1.0 generator on this site. Keep the references labeled the way you will label them after launch, so the brief does not have to be rewritten.
What you can actually run today
As of September 28, 2026:
- ByteDance’s published model post is still Seed Audio 1.0.
- Seed Audio 2.0 lists the new model as coming soon. The page offers early-bird credits for when that model is switched on, and the studio on the page still renders with 1.0 so a prompt can be tested first.
- The generator on seedaudio.co also runs Seed Audio 1.0. The Seed Audio 2.0 guide compares the published controls with the workspace you can use here.
Check the product page again before you promise a client six-minute stems or a video-aware dub. Those controls are specified. They are not what you get if the model menu still says 1.0.
Sources
- ByteDance Seed, From Speech to Audio Creation, July 20, 2026. Official description of Seed Audio 1.0 and the roadmap toward video input, longer scenes, and multitrack output.
- Seed Audio 2.0, reviewed September 28, 2026. Product spec for the six-minute window, six references, thirty languages, four input modes, stems, and the coming-soon status of the new model.
- Secondary explainers that restate that same brief, including scene-generation write-ups published in September 2026. Used only to see which claims are being repeated, not to add features the product page does not list.