All blog posts

Introducing the Hume Voice Controllability Leaderboard

HS
SM
Alice Baird
Hoon Shin, Shahbaz Mogal, and Alice Baird
··article
Share
Blog Image [whisper] Tag Applied

Modern text-to-speech models can generate natural, expressive speech. For narration, games, and interactive experiences, developers and other users of TTS systems also need the ability to direct who the voice sounds like, and the vocal characteristics and expressions of the spoken text. Great voice design models must be highly controllable.

Hume’s Real World VoiceEQ Benchmark evaluated 40+ voice models across 60+ metrics, with more than a million human ratings collected during its development. Today we are publishing the Voice Controllability Leaderboard, extending our evaluation approach to how well voice models follow direction.

How we evaluated voice controllability

We developed a multi-dimensional evaluation around the capabilities that matter for voice controllability: creating the requested voice, directing its delivery, following inline instructions and fitting the intended role. Our leaderboard covers 17 models, with different systems evaluated depending on the capability they support.

Several existing instruction-following TTS benchmarks use language models as judges. Our research shows these judgments do not always reflect human listeners’ experience especially given the possibility of data leakage: models that are iterated on with LLM judge predictions will likely have their scores inflated compared to human judgment. Thus, we opt for human ratings as a gold standard for the vast majority of our evaluations.

In total, the study draws on tens of thousands of human judgments, alongside objective measurements applied consistently within each evaluation.

Voice Design: Can the model build the voice a description asks for?

If the model receives the prompt for the voice of an authoritative British man in his 40s it should be able to hit each target reliably without losing clarity on its other directions.

For identity, raters identify the age, gender, and accent, and pick the correct voice from six voice texture pairs (e.g. raspy and gravelly vs. smooth and clear). For prosody, we measure pitch, tempo, volume, and pitch range acoustically against just-noticeable-difference thresholds.

We test American, British, Indian, and Australian accents, alongside regional US and UK varieties, world Englishes, and English influenced by other first languages. We also evaluate accents across 13 languages, with native raters scoring accent authenticity, native-speaker likeness, naturalness, and listenability.

Voice Instruction: Can it change how a fixed voice speaks?

Voice instruction evaluates the ability of a model to layer acting instructions on top of a preset voice. Like text-based LLMs, many TTS models use a system prompt to specify how the tone or emotion of a voice should change when reading a line, similar to a brief given to a professional voice actor.

For this eval, we take a similar approach as controllability, curating commonly requested dimensions of tone: authority, warmth, formality, and sarcasm alongside emotion and emotional intensity (i.e. does a furious voice sound stronger than a merely angry voice).

To reduce the impact of differences in individual raters’ standards, we used a forced-choice evaluation. Raters hear two samples generated by the same model with contrasting instructions and choose which better expresses the target attribute. For authority, for example, one sample is instructed to sound “tentative and deferential” and the other “commanding and authoritative.” Raters choose which sounds more authoritative, and the proportion of correct responses becomes the model’s score. This measures how clearly the model distinguishes the two ends of each axis.

Recent model releases (e.g. Elevenlabs v3, Inworld TTS 2, Higgs v3) opt for inline tags instead of prompt-wide instruction where users can providing information both on the placement AND type of delivery, enabling models to more easily align instructions with text.

We test 21 tags across 118 scripts, including vocal bursts such as [laughs], delivery styles such as [whispers], and prosodic instructions such as [very fast]. Human raters score execution, placement, and naturalness from 1 to 5, including how well models switch between sequential directions.

Role Fit: Which voice would you cast for a particular brief?

A model may reliably produce a voice that matches every detail in the description, yet still sound unconvincing. How those qualities come together – the gestalt of the voice – is important to believability and trust.

And in practice, users rarely prompt with voice characteristics in isolation, instead combining these qualities with specific use cases or character archetypes. To evaluate performance more holistically, we assess models for role fit, comparing models head-to-head on rich text prompts collected from our own production data.

Raters compare two voices generated from the same real production prompt, choosing which better matches the description and which sounds more natural. The prompts span 16 use-case categories derived from actual usage, including narration, characters, and customer service.

Results

The current results show no single model dominates every form of controllability. Google and ElevenLabs lead the most categories overall, but each has distinct weaknesses. Open models are still competitive in voice design, but recent improvements by proprietary models across different dimensions, language and use cases are evident.

Voice Design
Can a model build the requested voice from a text description alone? Each attribute is scored on its own terms — some by paid blind raters, prosody by direct measurement. Chance and units differ by tab, so read each caption before comparing. Best value per column in accent color; hover a score for detail.
Accuracy pooled over all six accent tasks. Chance varies from 17% to 33% by task, so columns are not comparable to one another.
#ProviderOverall ▼BasicUK
Regional
US
Regional
European
English
Asian & ME
English
World
Englishes
1elevenlabs/eleven_ttv_v3
proprietary
45.469.723.428.061.771.638.0
2inworld/voice-design
proprietary
35.873.639.235.64.915.329.9
3openbmb/VoxCPM2
open-source
32.861.321.620.414.236.349.1
4fish-speech/s2.1-pro
proprietary
25.745.612.528.216.724.132.7
5elevenlabs/eleven_multilingual_ttv_v2
proprietary
23.253.917.022.77.724.110.2
6qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign
open-source
10.125.03.722.02.52.52.2
Ranked by overall accuracy. Every model does best on the four basic accents and falls away outside them. The pooled leader does not lead the breakdown: Inworld wins three of the six families.hume.ai/rw-voice-eq

Across Voice Design, ElevenLabs v3 led in accent, age, texture, and multilingual performance, but ranked fourth of six on prosody. Models reliably conveyed the requested gender, averaging 97.87% accuracy, while age proved harder to control, with average accuracy of 51.8% and wider variation across models. Nasally-voices were one of the most difficult voice textures for models to produce (55.02%).

Performance varied sharply across accents and languages in ways an overall score hides. For example, even though Inworld’s Voice-Design led the most accented English test categories (3 out of 6), its poor showing European English and Asian and ME English meant elevenlabs/eleven_ttv_v3 had an higher overall score.

Voice design: French-accented English

Can the model create a voice speaking English with a convincing Parisian French accent that listeners can correctly identify and find easy to listen to?

Highly ratedeuro_french_good
0:00
Low ratedeuro_french_bad
0:00

ElevenLabs v3 topped seven of the 13 languages tested, with Inworld Voice Design leading in Chinese, German, Hindi, Japanese and Vietnamese with notable weak spots in Arabic, Indonesian, and Italian.2

Voice Instruction
Can a model change how a fixed voice speaks when asked? Each instruction is applied to the provider’s preset voice over a bank of short scripts, and scored by paid blind raters. Chance differs between the two tabs, so read each caption before comparing. Best value per column in accent color; hover a score for detail.
Raters hear two takes of the same voice and pick the one that follows the instruction, for example which sounds more commanding. Chance is 50%.
#ProviderOverall ▼AuthorityWarmthFormalitySarcasmIntensity
1google/gemini-3.1-flash
proprietary
90.295.292.990.979.288.0
2google/gemini-2.5-flash
proprietary
86.691.391.379.476.490.7
3google/gemini-2.5-pro
proprietary
82.888.991.370.680.681.5
4qwen/Qwen3-TTS-12Hz-1.7B
open-source
63.884.954.047.655.675.0
5openai/gpt-4o-mini-tts
proprietary
59.557.957.149.277.863.9
6bosonai/higgs-audio-v2
open-source
57.751.671.460.351.450.0
Ranked by overall accuracy. Sarcasm is the hardest instruction for the leaders and the one exception for gpt-4o-mini-tts, which is bottom two on every other tone but third on sarcasm. Two models sit at or near chance on individual tones.hume.ai/rw-voice-eq

Google’s family of models led in every tone and emotion category evaluated, whereas OpenAI’s gpt-4o-mini-tts struggled, particularly with emotion instructions where it ranked last.

Across the six models, authority was the most reliably conveyed tone, averaging 78.3% accuracy, while formality was the hardest at 66.3%. Sadness was the easiest emotion for listeners to identify, averaging 60.9%, while fear and disgust proved the most difficult at 30.8% and 30.0%, respectively.

Voice instruction: Fear

Can the model create a voice reading the line with fear that raters identify as fearful?

Highly ratedfear_good
0:00
Low ratedfear_bad
0:00

With style tags, our evaluation showed that the number of instructions and placement impacted performance. Overall, Inworld TTS 2 scored highest on vocal bursts (4.36/5), while Gemini 3.1 Flash led on delivery (4.26/5) and prosody (4.58/5). Gemini had an especially large lead on prosody with the next best model, Inworld, scoring 0.32 points below it.

However, across six of seven models scored lower when tags were combined than when tested individually with Gemini 3.1 Flash was the exception. Similarly for six of seven models, delivery tags at the start of a clip scored higher than those inserted mid-script. Vocal bursts at the very end were another failure point, with models sometimes omitting them entirely.

Role Fit
Which voice would you cast for the role? Paid blind raters hear two voices built for the same brief and vote for the better fit, and for the more natural one. All values are win rates over decided votes, so 50% is even. Best value in accent color; hover a score for detail.
Raters hear two voices built for the same production brief and vote for the better fit, and separately for the more natural one. Win rate over decided votes, so 50% is even rather than zero. Unreleased Hume systems are hidden, but matches against them still count.
#ProviderRole Fit ▼Naturalness
1elevenlabs/eleven_ttv_v3
proprietary
69.061.2
2google/gemini-3.1-flash
proprietary
66.956.1
3inworld/voice-design
proprietary
49.148.9
4elevenlabs/eleven_multilingual_ttv_v2
proprietary
44.753.1
5qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign
open-source
37.135.5
6openbmb/VoxCPM2
open-source
34.335.2
7fish-speech/s2.1-pro
proprietary
33.034.6
Ranked by role fit win rate, where 50% is even. The two columns come apart: eleven_multilingual_ttv_v2 loses on fit at 44.7% but wins on naturalness at 53.1%. Following the brief and sounding like a person are separable, and a model can buy one with the other.hume.ai/rw-voice-eq

Inworld led for podcast voices (90%) and therapist voices (64%), Gemini 3.1 Flash led for commercial (71%) and influencer (74%) voices, and ElevenLabs v3 led for educational (71%) and sports (90%) voices.

Voice fit: Therapist

Highly ratedtherapist_good
0:00
Low ratedtherapist_bad
0:00

Some models also showed especially large swings between use cases. Fish, for example, had a 33% win rate overall, but 82% for motivational voices and just 6% for meditation. Similarly, while Qwen was far behind closed-source models with a win rate of 37%, it was highly convincing on sports voices achieving a 67% win rate, second only to 11labs.

Explore the leaderboard

Explore the Voice Controllability Leaderboard on Hugging Face for per-model breakdowns and audio examples of successful and unsuccessful attempts.

Public leaderboards provide a shared baseline. Hume also works with frontier labs and product teams to design evaluations around their own voices, instructions, and production use cases, combining human judgments with objective measures. Get in touch to discuss evaluations tailored to your voice models and applications.



Stay in the loop

Get the latest on empathic AI research, product updates, and company news.

Join the community

Connect with other developers, share projects, and get help from the team.

Join our Discord