动察 Beating AI News Flash: ElevenLabs has released its next-generation speech models, Eleven v4 and the real-time version Eleven v4 Turbo. v4 focuses on more natural emotion and voice control, while Turbo is aimed at real-time voice agents, with a median inference latency of about 100ms.
On Artificial Analysis's latest TTS blind test leaderboard, Eleven v4 currently ranks first with 1315 Elo. Cartesia Sonic 3.6 is at 1275, Gemini 3.8 Flash TTS is at 1267, and Qwen-Audio-3.0-TTS-Plus is at 1258. The previous-generation Eleven v3 currently sits at only 1169, ranking 17th.
v3 already supported audio tags such as laughter and whispers, and v4 mainly makes this control more precise. Users can directly specify tone, pacing, emotion, and sound effects, and can also use natural language to describe how a sentence should be spoken. The model also expands language coverage from more than 70 to more than 90 languages, Instant Voice Clone requires only about 10 seconds of audio, and it enhances voice consistency in long-form text and multi-person conversations.
Turbo, meanwhile, is specifically designed to solve the speed problem in real-time conversation. Official documentation shows that v3 Conversational has a median inference latency of about 280ms, while v4 Turbo reduces it to about 100ms.

