Eleven v4 approaches synthetic speech as a performance

ElevenLabs is updating its speech generation stack with Eleven v4, a new architecture designed to follow performance directions, preserve speaker identity across regenerations, and maintain continuity in long-form audio. A Turbo variant targets voice agents with around 100 ms median inference latency.

Voice direction written directly into the script

A laugh, a whisper, or a door slam can now be part of the instructions given to the model. With Eleven v4, ElevenLabs introduces a new architecture designed to interpret the context of a script, identify different speakers, and adjust how each line is delivered.

Directions are written directly into the text using Audio Tags such as `[laughs]`, `[whispers]`, `[excited]`, or `[door slams]`. ElevenLabs says tag following is more reliable than with v3, including sequences combining multiple directions and sound effects.

The model supports more than 90 languages, multiple speakers within the same generation, and sound effects embedded into the scene. Regenerate a line without changing the character

Speaker stability is one of the main changes from Eleven v3. The same line can be regenerated multiple times while retaining the speaker's identity across dialogue and narration.

Professional Voice Clones, which were not supported in v3, also return with Eleven v4. They can use the new model's expressive and multilingual capabilities. ElevenLabs states that every voice clone requires verified consent from the voice owner.

For long-form content, context stitching carries context across multiple generations to reduce changes in pacing and delivery. A single generation remains limited to 10,000 characters, but multiple segments can be combined for formats such as audiobooks. A Turbo version for real-time conversations

Eleven v4 Turbo brings v4's expressive capabilities into a model designed for voice agents and live interactions.

ElevenLabs reports around 100 ms median inference latency and 150 ms median time to first speech. In the company's published comparison, Cartesia Sonic 3.6 reaches 262 ms time to first speech, while GPT-4o mini TTS reaches 814 ms.

Turbo supports bidirectional streaming. Text can be pushed progressively as an LLM generates it, while audio starts returning before the complete sentence is available.

Professional Voice Clones work across both v4 and v4 Turbo, allowing the same voice identity to move between produced content and real-time conversations. Clone, describe, or correct a voice

Eleven v4 retains several ways to create and control voices. A voice can be cloned from an audio sample, while Voice Design generates one from a written description.

Pronunciation dictionaries remain available for names, acronyms, and technical vocabulary. Pauses and performance directions now rely on natural-language Audio Tags, while SSML tags such as `` are disabled in v4.

ElevenLabs' library includes more than 17,500 voices compatible with the model. Existing Instant Voice Clones and Professional Voice Clones created before v4 need to be retrained to work effectively with the new architecture. The same expressive range for production and agents

Eleven v4 is primarily tuned for produced content where voice quality takes priority, while v4 Turbo is designed for interactive applications and ElevenAgents.

Both models are available through the API with streaming and non-streaming endpoints, alongside TypeScript and Python SDKs. Output formats remain consistent with other ElevenLabs models, including MP3, WAV, PCM, and µ-law for telephony.