跳到正文
原文
The Decoder· Jonathan Kemper·· 4 小时前AI 评分67

ElevenLabs发布Eleven v4语音模型,提升表现力与一致性

ElevenLabs' new v4 speech model makes AI voices more expressive and consistent

AI 导读

ElevenLabs发布语音模型Eleven v4,能更准确跟随脚本中的情绪、停顿和音效指令,并在长制作中保持音色一致;同一新架构还支撑面向实时语音智能体的Turbo变体。

正文

Elevenlabs is releasing Eleven v4, a new speech model that follows direction cues more accurately and keeps voices consistent across long productions. A new model architecture also powers the Turbo variant for real-time voice agents.

Eleven v4 generates laughter, whispers, and sounds like slamming doors more reliably than its predecessor. The Turbo variant for voice agents starts producing speech in about 150 milliseconds. Eleven v3, released just over a year ago, already supported these audio tags but followed them less accurately.

Elevenlabs interface showing a dialogue script for two speakers, with orange-highlighted tags like warm, long pause, Gong sounds, and nervous laugh in the text, along with settings for voice, model Eleven v4, stability, and similarity.
Tags in the script let users set emotions, pauses, and sound effects for each line. | Image: Elevenlabs

More consistency

Elevenlabs says v4 uses a new architecture that analyzes a script's tone, pacing, and context. Users can give directions through tags or plain sentences and use phonetic spelling to set the pronunciation of names and technical terms. The company says pronunciation controls now work more reliably. Narrators and characters should sound consistent throughout a production, even when users regenerate individual lines several times.

Eleven v4 handles up to 10,000 characters per request, roughly ten minutes of audio. Longer works like audiobooks use multiple segments, with pacing and delivery expected to stay consistent across transitions. In dialogue, AI speakers respond to the context of the entire scene rather than delivering each line in isolation.

The new version supports more than 90 languages, up from about 70 with v3. Cloned voices should speak other languages with native accents without drifting back to their original accents over time. Professional Voice Clones work again after being unsupported in v3, while an "Instant Voice Clone" needs only ten seconds of audio.

Turbo aims to make expressive voice agents faster

Elevenlabs is also releasing a faster v4 variant for real-time uses like customer service calls or game characters. The company says voice agent developers previously had to choose between speed and expression, and v4 Turbo is meant to offer both.

In Elevenlabs' tests, Turbo starts producing audible speech in 150 milliseconds, compared with 262 milliseconds for Cartesia Sonic 3.6. OpenAI's GPT-4o mini TTS takes 814 milliseconds. Elevenlabs optimized Turbo together with its ElevenAgents platform.

Bar chart showing median time to first audible speech, with Eleven v4 Turbo at 150 ms ahead of Cartesia Sonic 3.6 at 262 ms, xAI TTS at 362 ms, Gemini 3.8 Flash-Lite TTS at 685 ms, and OpenAI GPT-4o mini TTS at 814 ms.
In Elevenlabs' benchmarks, Eleven v4 Turbo starts producing speech much faster than the competing models tested. | Image: Elevenlabs

Eleven v4 ranks ahead of Cartesia Sonic 3.6 and Google's Gemini 3.8 Flash TTS on Artificial Analysis' Provider Voice Arena leaderboard. It scores 91.7 percent on the pronunciation benchmark, up from v3's 85.6 percent. In Elevenlabs' blind tests, about three-quarters of listeners preferred v4 over models from Cartesia, Inworld, and Google.

Bar chart showing the share of blind head-to-head comparisons won by Eleven v4, with 81 percent against Cartesia Sonic 3.6 and Inworld TTS-2, 72 percent against Gemini 3.8 Flash-Lite TTS, and 65 percent against Gemini 3.8 Flash TTS.
In the company's blind tests, listeners rated Eleven v4 as more expressive in 65 to 81 percent of comparisons, depending on the competitor. | Image: Elevenlabs

Higher-quality voice clones make v4 useful for dubbing, with Elevenlabs promoting the ability to use an actor's voice across all supported languages. The company licenses voices from the people behind them. Voice actors can offer an extensively trained clone of their voice in Elevenlabs' library and earn money when paying users use it. Elevenlabs already offers access to celebrity voices like Michael Caine's through a dedicated marketplace.

Both models launch with temporary price cuts

Elevenlabs' standard API pricing is $80 per million characters for v4 and $40 for Turbo. Through October 12, those rates fall to $22 and $11. According to Elevenlabs, users on the $22 monthly Creator plan or higher can use v4 in ElevenCreative at no extra cost for two weeks. Usage is capped at twice their monthly credits. Artificial Analysis lists Sonic 3.6 at $49 per million characters and Gemini 3.8 Flash TTS at $16.49.

Both models are available now in ElevenAgents, ElevenCreative, and through the API. Elevenlabs stores customer data in the US by default, according to its documentation. Enterprise customers can store data in isolated environments in the EU, India, or Singapore, though some processing may take place outside the chosen region. In the EU, customers can keep API processing within the region by using a mode that doesn't retain data.

Elevenlabs also released its Music 2.5 model in mid-September, designed to produce denser, more natural-sounding songs.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

来源:The Decoder · the-decoder.com