ElevenLabs Launches Eleven v4 and Eleven v4 Turbo Text-to-Speech Models
  • News
  • Europe

ElevenLabs Launches Eleven v4 and Eleven v4 Turbo Text-to-Speech Models

The new emotive models support over 90 languages and low-latency voice agents

9/28/2026
•Ghita Khalfaoui
Back to News

ElevenLabs has launched Eleven v4, its most emotive text-to-speech model to date, alongside a low-latency companion called Eleven v4 Turbo. The company said the release addresses a longstanding gap between rapid response times and natural, emotionally aware speech generation. Both models are available now through ElevenAgents, ElevenCreative, and the ElevenAPI, with a free account option for new users.


A New Architecture for Expressive Speech

Eleven v4 is built on an entirely new architecture and has been ranked first by Artificial Analysis among text-to-speech models. In blind head-to-head tests, listeners preferred it over competing models around 75% of the time. The system interprets tone, pacing, emotion, character, and context to produce speech that can sound dramatic, tender, urgent, comedic, or conversational while preserving speaker identity.

Fine-Grained Direction and Audio Tags

Users can describe how a line should be delivered in natural language and add instructions for specific phrases, emotions, or sound effects. Inline tags such as [laughs], [said angrily in French accent], [light rain], and [phone buzzing] are followed more accurately than in prior models. The company also improved support for International Phonetic Alphabet phonemes, making custom pronunciations more reliable.

Conversational Consistency Across Speakers

Eleven v4 introduces a new method for capturing speaker identities that preserves the unique qualities of each voice. The model understands the context of an entire scene, so multiple speakers respond to what has just been said rather than delivering isolated lines. This supports more natural dialogue in agent conversations, audiobooks, and advertisements.

A Turbo Model for Low-Latency Agents

Eleven v4 Turbo brings the same expressive technology to use cases that require immediate responses, with a median inference latency around 100 milliseconds and a median time to first speech around 150 milliseconds. These speeds are faster than the average pause between two people talking, the company noted. The model is designed for industries where fast and emotional voice agents are valuable, including healthcare and gaming.

Optimized for ElevenAgents

Eleven v4 Turbo was developed alongside the ElevenAgents conversational agents platform to ensure the models work as one system. This integration gives users more expressive, reliable, and low latency agents without the need to stitch together separate vendors. ElevenLabs said its research and engineering teams optimized both components together for individual use cases.

Multilingual Expression and Native Accents

Both models support more than 90 languages and capture rhythm, emotion, and delivery more effectively across linguistic contexts. A voice recorded in one language can speak another fluently while adopting a native accent and retaining the original identity. Accent adherence is stronger than before, reducing the tendency to drift back toward the source accent over longer generations.

Voice Cloning and Long-Form Consistency

Voice cloning is more authentic and consistent across both models, with significantly better speaker similarity to the original source. Instant Voice Clones can now capture voices with high fidelity using just 10 seconds of audio. Professional Voice Clones are also supported for the highest fidelity applications, while request stitching has become more reliable for long-form content.


ElevenLabs positions Eleven v4 and Eleven v4 Turbo as the culmination of its latest research in expressive speech generation, intended for content where delivery matters as much as the words themselves. The company is targeting audiobooks, character performances, voiceovers, dubbing, and localized conversational agents. Creators can create a free account to begin generating with either model immediately.