Article

ElevenLabs reaches the top spot in v4 TTS rankings, surpassing Gemini and Qwen

Beating AI News Flash:ElevenLabs has released the next-generation voice models Eleven v4 and the real-time Eleven v4 Turbo. v4 focuses on more natural emotion and voice control, while Turbo targets real-time voice Agents, with a median inference latency of approximately ms.

In the latest Artificial Analysis TTS blind test rankings, Eleven v4 currently ranks first with 1315 Elo. Cartesia Sonic 3.6 ranks 1275,Gemini 3.8 Flash TTS ranks 1267,Qwen-Audio-3.0-TTS-Plus ranks 1258. The previous-generation Eleven v3 currently has only 1169, ranking 17.

v3 already supports audio tags such as laughter and whispers; v4 mainly makes this control more precise. Users can directly specify tone, pacing, emotion, and sound effects, or describe in natural language how a sentence should be spoken. The model has also expanded language coverage from 70 languages to 90 languages, Instant Voice Clone requires only approximately 10 seconds of audio, and consistency of voices in long-form text and multi-speaker conversations has been improved.

Turbo specifically addresses the speed of real-time conversations. Official documentation shows that the median inference latency of v3 Conversational is approximately ms, while v4 Turbo reduces it to approximately ms.

Original link https://m.theblockbeats.info/flash/369554