ByteDance brings together voices, music, and sound effects in Seed Audio 1.0

ByteDance has launched Seed Audio 1.0, an AI model that generates speech, background music, and sound effects in a single pass via Volcano Engine and Doubao.

ByteDance is pushing audio synthesis further with Seed Audio 1.0, a model that goes beyond just voice. Developed by the Seed/Doubao team, it generates speech, background music, and sound effects in a single pass, without separate editing steps. The system supports multi-character dialogue with distinct voices, managing emotions, accents, and non-verbal elements. Rendering can be based on text input and guided by up to three reference audio clips or an image, with the user retaining control over speed, volume, pitch, and output format.

Technically, it's a non-streaming model: it produces the entire result in one block, up to two minutes per generation, maintaining voice consistency from one sequence to the next.

Access is currently by invitation only. The model is available via the Volcano Engine Ark platform API and the Doubao application, while BytePlus, the international arm of the developer, is opening applications for enterprise access. On Chinese platforms, the tool is referenced as Doubao-Seed-Audio 1.0.