Back
ElevenLabs' Eleven v4 speech model makes AI voices more expressive and consistent
SiTech AI Team3 წთ. საკითხავი

ElevenLabs' Eleven v4 speech model makes AI voices more expressive and consistent

ElevenLabs has released Eleven v4, a speech model that follows direction cues more accurately and keeps voices stable across long productions. A Turbo variant for real-time voice agents responds in about 150 milliseconds.

ElevenLabs has released Eleven v4, a new speech model that follows direction cues more accurately and keeps a voice consistent across long productions. Eleven v4 Turbo, the variant built for real-time voice agents, starts producing speech in roughly 150 milliseconds.

More consistency

Eleven v3 shipped a little over a year ago and already handled laughter, whispers and effects such as slamming doors, but followed the cues less accurately. According to ElevenLabs, v4 uses a new architecture that analyzes a script's tone, pacing and context. Users can give directions through tags or plain sentences, and set the pronunciation of names and technical terms with phonetic spelling. The company says pronunciation controls now work more reliably.

The Eleven v4 interface with a script and tags

A single request can carry up to 10,000 characters, roughly ten minutes of audio, while longer works such as audiobooks are split into segments. The model supports more than 90 languages, up from about 70 in v3. Professional Voice Clones work again after going unsupported in v3, and an Instant Voice Clone needs only ten seconds of audio. Cloned voices should speak other languages with native accents without drifting back to their original accent over time.

Turbo targets real-time voice agents

ElevenLabs is also shipping a faster variant for customer service calls and game characters. The company says developers previously had to choose between speed and expression, and Turbo is meant to offer both. In ElevenLabs' tests it starts producing audible speech in 150 milliseconds, compared with 262 milliseconds for Cartesia Sonic 3.6 and 814 milliseconds for OpenAI's GPT-4o mini TTS. Turbo was optimized alongside the ElevenAgents platform.

Benchmarks and launch pricing

On Artificial Analysis' Provider Voice Arena leaderboard, Eleven v4 ranks ahead of Cartesia Sonic 3.6 and Google's Gemini 3.8 Flash TTS. It scores 91.7 percent on the pronunciation benchmark, up from v3's 85.6 percent. In the company's blind tests, about three quarters of listeners rated v4 as more expressive: it won 81 percent of comparisons against Cartesia Sonic 3.6 and Inworld TTS-2, 72 percent against Gemini 3.8 Flash-Lite TTS and 65 percent against Gemini 3.8 Flash TTS.

Standard API pricing is $80 per million characters for v4 and $40 for Turbo, falling to $22 and $11 through October 12. Users on the $22 monthly Creator plan or higher can use v4 in ElevenCreative at no extra cost for two weeks. Both models are available now in ElevenAgents, ElevenCreative and through the API.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.