![Blog Image [whisper] Tag Applied](/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2Fxqnc2for%2Fproduction%2F06414e1fae798399961068290af810114be8148e-1200x800.jpg%3Frect%3D0%2C63%2C1200%2C675%26w%3D2048%26h%3D1152&w=3840&q=75)
Modern text-to-speech models can generate natural, expressive speech. For narration, games, and interactive experiences, developers and other users of TTS systems also need the ability to direct who the voice sounds like, and the vocal characteristics and expressions of the spoken text. Great voice design models must be highly controllable.
Hume’s Real World VoiceEQ Benchmark evaluated 40+ voice models across 60+ metrics, with more than a million human ratings collected during its development. Today we are publishing the Voice Controllability Leaderboard, extending our evaluation approach to how well voice models follow direction.
How we evaluated voice controllability
We developed a multi-dimensional evaluation around the capabilities that matter for voice controllability: creating the requested voice, directing its delivery, following inline instructions and fitting the intended role. Our leaderboard covers 17 models, with different systems evaluated depending on the capability they support.
Several existing instruction-following TTS benchmarks use language models as judges. Our research shows these judgments do not always reflect human listeners’ experience especially given the possibility of data leakage: models that are iterated on with LLM judge predictions will likely have their scores inflated compared to human judgment. Thus, we opt for human ratings as a gold standard for the vast majority of our evaluations.
In total, the study draws on tens of thousands of human judgments, alongside objective measurements applied consistently within each evaluation.
Voice Design: Can the model build the voice a description asks for?
If the model receives the prompt for the voice of an authoritative British man in his 40s it should be able to hit each target reliably without losing clarity on its other directions.
For identity, raters identify the age, gender, and accent, and pick the correct voice from six voice texture pairs (e.g. raspy and gravelly vs. smooth and clear). For prosody, we measure pitch, tempo, volume, and pitch range acoustically against just-noticeable-difference thresholds.
We test American, British, Indian, and Australian accents, alongside regional US and UK varieties, world Englishes, and English influenced by other first languages. We also evaluate accents across 13 languages, with native raters scoring accent authenticity, native-speaker likeness, naturalness, and listenability.
Voice Instruction: Can it change how a fixed voice speaks?
Voice instruction evaluates the ability of a model to layer acting instructions on top of a preset voice. Like text-based LLMs, many TTS models use a system prompt to specify how the tone or emotion of a voice should change when reading a line, similar to a brief given to a professional voice actor.
For this eval, we take a similar approach as controllability, curating commonly requested dimensions of tone: authority, warmth, formality, and sarcasm alongside emotion and emotional intensity (i.e. does a furious voice sound stronger than a merely angry voice).
To reduce the impact of differences in individual raters’ standards, we used a forced-choice evaluation. Raters hear two samples generated by the same model with contrasting instructions and choose which better expresses the target attribute. For authority, for example, one sample is instructed to sound “tentative and deferential” and the other “commanding and authoritative.” Raters choose which sounds more authoritative, and the proportion of correct responses becomes the model’s score. This measures how clearly the model distinguishes the two ends of each axis.
Recent model releases (e.g. Elevenlabs v3, Inworld TTS 2, Higgs v3) opt for inline tags instead of prompt-wide instruction where users can providing information both on the placement AND type of delivery, enabling models to more easily align instructions with text.
We test 21 tags across 118 scripts, including vocal bursts such as [laughs], delivery styles such as [whispers], and prosodic instructions such as [very fast]. Human raters score execution, placement, and naturalness from 1 to 5, including how well models switch between sequential directions.
Role Fit: Which voice would you cast for a particular brief?
A model may reliably produce a voice that matches every detail in the description, yet still sound unconvincing. How those qualities come together – the gestalt of the voice – is important to believability and trust.
And in practice, users rarely prompt with voice characteristics in isolation, instead combining these qualities with specific use cases or character archetypes. To evaluate performance more holistically, we assess models for role fit, comparing models head-to-head on rich text prompts collected from our own production data.
Raters compare two voices generated from the same real production prompt, choosing which better matches the description and which sounds more natural. The prompts span 16 use-case categories derived from actual usage, including narration, characters, and customer service.
Results
The current results show no single model dominates every form of controllability. Google and ElevenLabs lead the most categories overall, but each has distinct weaknesses. Open models are still competitive in voice design, but recent improvements by proprietary models across different dimensions, language and use cases are evident.
| # | Provider | Overall ▼ | Basic | UK Regional | US Regional | European English | Asian & ME English | World Englishes |
|---|---|---|---|---|---|---|---|---|
| 1 | elevenlabs/eleven_ttv_v3 proprietary | 45.4 | 69.7 | 23.4 | 28.0 | 61.7 | 71.6 | 38.0 |
| 2 | inworld/voice-design proprietary | 35.8 | 73.6 | 39.2 | 35.6 | 4.9 | 15.3 | 29.9 |
| 3 | openbmb/VoxCPM2 open-source | 32.8 | 61.3 | 21.6 | 20.4 | 14.2 | 36.3 | 49.1 |
| 4 | fish-speech/s2.1-pro proprietary | 25.7 | 45.6 | 12.5 | 28.2 | 16.7 | 24.1 | 32.7 |
| 5 | elevenlabs/eleven_multilingual_ttv_v2 proprietary | 23.2 | 53.9 | 17.0 | 22.7 | 7.7 | 24.1 | 10.2 |
| 6 | qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign open-source | 10.1 | 25.0 | 3.7 | 22.0 | 2.5 | 2.5 | 2.2 |
Across Voice Design, ElevenLabs v3 led in accent, age, texture, and multilingual performance, but ranked fourth of six on prosody. Models reliably conveyed the requested gender, averaging 97.87% accuracy, while age proved harder to control, with average accuracy of 51.8% and wider variation across models. Nasally-voices were one of the most difficult voice textures for models to produce (55.02%).
Performance varied sharply across accents and languages in ways an overall score hides. For example, even though Inworld’s Voice-Design led the most accented English test categories (3 out of 6), its poor showing European English and Asian and ME English meant elevenlabs/eleven_ttv_v3 had an higher overall score.
Voice design: French-accented English
Can the model create a voice speaking English with a convincing Parisian French accent that listeners can correctly identify and find easy to listen to?
ElevenLabs v3 topped seven of the 13 languages tested, with Inworld Voice Design leading in Chinese, German, Hindi, Japanese and Vietnamese with notable weak spots in Arabic, Indonesian, and Italian.2
| # | Provider | Overall ▼ | Authority | Warmth | Formality | Sarcasm | Intensity |
|---|---|---|---|---|---|---|---|
| 1 | google/gemini-3.1-flash proprietary | 90.2 | 95.2 | 92.9 | 90.9 | 79.2 | 88.0 |
| 2 | google/gemini-2.5-flash proprietary | 86.6 | 91.3 | 91.3 | 79.4 | 76.4 | 90.7 |
| 3 | google/gemini-2.5-pro proprietary | 82.8 | 88.9 | 91.3 | 70.6 | 80.6 | 81.5 |
| 4 | qwen/Qwen3-TTS-12Hz-1.7B open-source | 63.8 | 84.9 | 54.0 | 47.6 | 55.6 | 75.0 |
| 5 | openai/gpt-4o-mini-tts proprietary | 59.5 | 57.9 | 57.1 | 49.2 | 77.8 | 63.9 |
| 6 | bosonai/higgs-audio-v2 open-source | 57.7 | 51.6 | 71.4 | 60.3 | 51.4 | 50.0 |
Google’s family of models led in every tone and emotion category evaluated, whereas OpenAI’s gpt-4o-mini-tts struggled, particularly with emotion instructions where it ranked last.
Across the six models, authority was the most reliably conveyed tone, averaging 78.3% accuracy, while formality was the hardest at 66.3%. Sadness was the easiest emotion for listeners to identify, averaging 60.9%, while fear and disgust proved the most difficult at 30.8% and 30.0%, respectively.
Voice instruction: Fear
Can the model create a voice reading the line with fear that raters identify as fearful?
With style tags, our evaluation showed that the number of instructions and placement impacted performance. Overall, Inworld TTS 2 scored highest on vocal bursts (4.36/5), while Gemini 3.1 Flash led on delivery (4.26/5) and prosody (4.58/5). Gemini had an especially large lead on prosody with the next best model, Inworld, scoring 0.32 points below it.
However, across six of seven models scored lower when tags were combined than when tested individually with Gemini 3.1 Flash was the exception. Similarly for six of seven models, delivery tags at the start of a clip scored higher than those inserted mid-script. Vocal bursts at the very end were another failure point, with models sometimes omitting them entirely.
| # | Provider | Role Fit ▼ | Naturalness |
|---|---|---|---|
| 1 | elevenlabs/eleven_ttv_v3 proprietary | 69.0 | 61.2 |
| 2 | google/gemini-3.1-flash proprietary | 66.9 | 56.1 |
| 3 | inworld/voice-design proprietary | 49.1 | 48.9 |
| 4 | elevenlabs/eleven_multilingual_ttv_v2 proprietary | 44.7 | 53.1 |
| 5 | qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign open-source | 37.1 | 35.5 |
| 6 | openbmb/VoxCPM2 open-source | 34.3 | 35.2 |
| 7 | fish-speech/s2.1-pro proprietary | 33.0 | 34.6 |
Inworld led for podcast voices (90%) and therapist voices (64%), Gemini 3.1 Flash led for commercial (71%) and influencer (74%) voices, and ElevenLabs v3 led for educational (71%) and sports (90%) voices.
Voice fit: Therapist
Some models also showed especially large swings between use cases. Fish, for example, had a 33% win rate overall, but 82% for motivational voices and just 6% for meditation. Similarly, while Qwen was far behind closed-source models with a win rate of 37%, it was highly convincing on sports voices achieving a 67% win rate, second only to 11labs.
Explore the leaderboard
Explore the Voice Controllability Leaderboard on Hugging Face for per-model breakdowns and audio examples of successful and unsuccessful attempts.
Public leaderboards provide a shared baseline. Hume also works with frontier labs and product teams to design evaluations around their own voices, instructions, and production use cases, combining human judgments with objective measures. Get in touch to discuss evaluations tailored to your voice models and applications.



