Seed Audio 1.0
Sign up
Cinematic audio studio with layered voice, ambience, effects, and music waveforms converging into one sound scene
Independent capability guide · Updated August 2026

Seed Audio 2.0: complete audio creation from text, voice, and video

Seed Audio 2.0 is presented as a unified workflow for creating dialogue, emotional performance, ambience, foley, sound effects, and music from one connected brief. This guide explains the published T2A, TA2A, TV2A, and TAV2A modes, what changed from the 1.0 baseline, and what is actually available on this site today.

Published input modes

T2A, TA2A, TV2A, and TAV2A

Reference capacity

Up to six reference audios on the Dreamina workflow

Published output length

Up to six minutes per result

Generator on this page

Current Seed Audio 1.0 integration

Start generating Seed Audio 1.0 online.

Use the Seed Audio 1.0 workspace below to create sound scenes from a prompt, optional reference audio, or one reference image.

Input

Prompt-first audio generation with optional controls.

Prompt*
83/2048

Additional Settings

Customize your input with more control.

History

Your recent Seed Audio 1.0 Preview generations.

Sign in to see your generation history.

Model direction

Direct the whole sound scene, not a pile of isolated clips

The practical idea behind Seed Audio 2.0 is that a creator should be able to describe the relationship between speech, setting, action, music, and timing before generation begins. A tense line should sit inside the same room tone as the footsteps that interrupt it, while a music cue should rise at the intended dramatic beat.

Dreamina describes a workflow that keeps those layers connected through text prompts, permitted reference voices, optional reference video, timestamps, and separate tracks. That makes seedaudio 2.0 relevant to short films, dubbing, podcasts, audiobooks, advertising, games, animation, and other projects where sound must follow a scene rather than merely read a script.

New Seed Audio 2.0 concept image showing dialogue, ambience, foley, and music layers forming a rainy cinematic scene
Input map

Four paths into one reviewable audio result

Seed Audio 2.0 is described through four mode names. The useful difference is not the acronym itself, but which source material the model can use to preserve voice identity, follow visual timing, and shape the complete mix.

T2A · Text to audio

Start with a written production brief. Define speakers, dialogue, delivery, environment, foley events, sound effects, music direction, pacing, and the desired ending so the result is composed as one scene.

TA2A · Reference audio control

Add a permitted voice reference when identity matters. Describe which dimensions may change, such as emotion, rhythm, accent, intensity, pauses, or non-speech expression, while keeping the speaker recognizable.

TV2A · Video-aware audio

Supply a reference video plus a text direction. The published workflow uses visual cuts and action beats to guide voiceover, atmosphere, effects, music, and placement across the edit.

TAV2A · Voice plus video

Combine a permitted voice reference with video context and written direction. This path is intended for character-consistent dubbing or scoring where both vocal identity and visual timing matter.

1.0 baseline versus published 2.0 direction

The clearest changes are references, duration, video context, and editability

ByteDance Seed's official 1.0 announcement provides the technical baseline. Dreamina's Seed Audio 2.0 page describes the expanded product direction below; public API details and independent benchmarks are not yet assumed.

Workflow area
Seed Audio 1.0 baseline
Published 2.0 direction
Complete scenes
Unified speech, emotion, timing, ambience, and sound effects from one scene prompt.
Keeps dialogue, ambience, foley, effects, and music connected across a longer production workflow.
References
The official model supports authorized reference voice guidance; this site's current tool accepts up to three audio inputs in total.
Dreamina states support for up to six reference audios, with a separate speaker represented by each reference file.
Duration
ByteDance Seed documents up to two minutes in one pass, with continuation for longer material.
Dreamina states that one generated result can run for up to six minutes.
Video context
Video references were described as a future direction in the official 1.0 launch note.
TV2A and TAV2A are presented as video-aware paths for dubbing, atmosphere, effects, and scoring.
Timing and tracks
Prompt-level dialogue timing is documented at 100 ms intervals; other sound layers have less granular control.
Timestamps, separate tracks, and reusable voice assets are presented as tools for review and mix revisions.
Prompt and review workflow

Write a production brief that survives the first generation

A strong Seed Audio 2.0 request explains relationships and timing instead of listing disconnected keywords. Build the prompt in passes so failed outputs are easier to diagnose and the next revision changes only one creative variable.

New concept image of video-aware dubbing with voice, ambience, effects, and music arranged on separate synchronized waveform tracks

Step 1

Define the dramatic job

State the scene type, duration target, language, speakers, emotional turn, location, and final beat. This gives every later sound layer a shared narrative purpose.

Step 2

Map references to speakers

Use only voices you are allowed to use, keep one speaker per reference file, and describe what may change without asking the model to imitate an identifiable person without consent.

Step 3

Place dialogue and events

Name the order of lines, pauses, actions, ambience changes, and music entries. When video is supplied, connect those instructions to visible cuts or action beats instead of relying on a generic mood.

Step 4

Review layer by layer

Check voice identity, intelligibility, timing, room continuity, effects, and music separately. Revise the weakest dimension first so a useful performance is not discarded because one background layer needs work.

Example complete-scene direction

Create a 45-second original radio-drama scene in a nearly empty railway station at night. Two fictional speakers exchange restrained dialogue in English. Keep their voices distinct, place a distant train before the second line, add wet footsteps and a quiet electrical hum, then bring in a low original music bed for the final ten seconds. Leave brief breathing space between lines and end on the station announcement fading into rain.

Availability and responsible use

Know which claims belong to the model, the product page, and this site

Seed Audio 2.0 is attracting search interest before a complete public technical and API record is easy to verify. These boundaries help teams evaluate the workflow without turning marketing descriptions into unsupported guarantees.

Dreamina publishes the named 2.0 workflow, while ByteDance Seed's public model directory currently lists Seed Audio 1.0. Treat public API availability, pricing, release dates, and benchmarks as unconfirmed until an official source documents them.

Reference voices, videos, scripts, music directions, and final outputs may involve copyright, privacy, publicity, or contractual rights. Obtain consent and review the terms of the access channel used for generation.

Longer, layered audio can still produce timing conflicts, masking, inconsistent identity, or unwanted artifacts. Human review and selective regeneration remain part of a production workflow.

Frequently asked questions

Practical answers before you choose a mode or write a prompt

These answers separate the published 2.0 direction from the capabilities of the current generator on seedaudio.co.

What is Seed Audio 2.0?

Seed Audio 2.0 is presented by Dreamina as a unified audio-generation workflow for dialogue, emotional voice performance, ambience, foley, effects, music, reference-audio control, and video-aware dubbing. This page treats that product description as the source and does not infer undocumented API behavior.

What do T2A, TA2A, TV2A, and TAV2A mean?

T2A starts from text, TA2A adds reference audio, TV2A adds reference video, and TAV2A combines text, permitted reference audio, and video. The added context determines whether the workflow can focus on scene composition, voice identity, visual timing, or all three.

How is Seed Audio 2.0 different from Seed Audio 1.0?

The published differences center on up to six reference audios, output up to six minutes, video-aware TV2A and TAV2A paths, timestamps, separate tracks, and reusable voice assets. The full-scene idea itself already exists in the official 1.0 model.

How long can it generate and how many references can it use?

Dreamina states up to six minutes of output and up to six reference audios for the 2.0 workflow. Those limits do not apply to the generator embedded here, which currently follows this site's Seed Audio 1.0 constraints.

Is the generator on this page running Seed Audio 2.0?

No. It is the same production Seed Audio 1.0 generator used on the homepage. It is provided so you can practice complete-scene prompting while keeping the model version explicit.

Is there a public seed-audio-2.0 API on seedaudio.co?

Not currently. The site's public audio API and web generator use the existing 1.0 integration. A future 2.0 integration would require verified provider documentation, request fields, billing rules, callbacks, and result handling before it could be advertised.

Can seedaudio 2.0 output be used commercially?

Commercial use depends on the terms of the platform that actually generates the audio and on the rights attached to every prompt, voice, video, script, and output. Review those terms and obtain permission rather than assuming a universal license.

Research trail

Read the published claims and the 1.0 technical baseline

The factual core of this Seed Audio 2.0 guide comes from Dreamina's model page and ByteDance Seed's official Seed Audio 1.0 launch article. X search was attempted through OpenCLI but returned a session rate limit, so no social claim is presented as evidence.