Real World VoiceEQ
Voice AI has made remarkable progress.
Faster responses.Fewer errors.
Something still feels off.
Scroll
TTS leaderboard
Text-to-speech, scored on what it produces
We score how well a model turns text into speech across the dimensions that decide real-world quality: expressive range, identity stability, pronunciation accuracy, and audio cleanliness. What matters is what the model produces across a full performance, not how clean one sentence sounds.
Overall
Frontier score: closeness to the ideal corner (TOPSIS) on two rater-controlled axes, Expressivity and Reliability, each built from the individual rater questions that measure it across all evals. Models without an instruction channel read the same prompts without the direction. 0-1; higher is better.
- 01
google/gemini-3.8-flashProprietary0.92
- 02
google/gemini-3.8-flash-liteProprietary0.91
- 03
google/gemini-2.5-proProprietary0.88
- 04
google/gemini-2.5-flashProprietary0.86
- 05
cartesia/sonic-3.6Proprietary0.84
- 06
xai/grok-ttsProprietary0.79
- 07
google/gemini-3.1-flashProprietary0.78
- 08
elevenlabs/eleven_v3_conversationalProprietary0.77
- 09
cartesia/sonic-3.5Proprietary0.76
- 10speechify/simba-3.2Proprietary0.75
- 11
openai/gpt-4o-mini-ttsProprietary0.74
- 12
fishaudio/s2-proOpen0.72
- 13
elevenlabs/eleven_v3Proprietary0.71
- 14
openai/tts-1Proprietary0.69
- 15
openai/tts-1-hdProprietary0.69
- 16
inworld/tts-2-flashProprietary0.68
- 17
inworld/tts-1-maxProprietary0.65
- 18
elevenlabs/eleven_multilingual_v2Proprietary0.64
- 19
deepgram/aura-2Proprietary0.59
- 20smallest/lightning_v3.1Proprietary0.58
- 21
inworld/tts-2Proprietary0.58
- 22
elevenlabs/eleven_flash_v2_5Proprietary0.57
- 23
microsoft/VibeVoice-1.5BOpen0.56
- 24hexgrad/Kokoro-82MOpen0.54
- 25
inworld/tts-1Proprietary0.54
- 26
resemble-ai/chatterboxOpen0.51
- 27
qwen/Qwen3-TTS-12Hz-1.7BOpen0.49
- 28
hume/tada-3b-mlOpen0.48
- 29sesame/csm-1bOpen0.38
- 30
coqui/xtts-v2Open0.37
- 31bosonai/higgs-audio-v2Open0.35
- 32bosonai/higgs-audio-v3-tts-4bOpen0.29
- 33
indexteam/indextts-2Open0.18
- 34nari-labs/diaOpen0.17
- 35
parler-tts/parler-ttsOpen0.04
Explore the complete results.
View full model rankings, audio samples, methodology, per-dimension results, and additional benchmark analysis.
Private evaluations
Evaluate your system in the conditions that matter.
A public benchmark shows how you compare at a high level. A private evaluation gives you specific data and insights into where your system falls short and what to improve.
Built on the same research and evaluation infrastructure, tailored to your system and use case.
Your system
Model, agent, endpoint, or checkpoint
Your criteria
Tailored to the modality and use case
Your conditions
Scenarios, speakers, and environments relevant to your users
Your results, private
Condition-level analysis, kept to your team
Checkpoint comparison
Internal checkpoints or external systems, side by side
Reusable regression
A held-out suite to rerun on the next release
How we measure
The right judge for every measure.
Different qualities require different ways of measuring them. Hume matches each metric to objective measures, model graders, or trained human raters, combining approaches where needed.
- 01
Objective measures
- Accuracy
- Timing
- Acoustics
- 02
Automated graders
- Defined criteria
- At scale
- Checked against human ratings
- 03
Human raters
- Naturalness
- Expression
- Context
What a private evaluation shows
Get the diagnostic-level detail to guide what you do next
A private evaluation returns capability-level results for your own system, and the conditions underneath each average.
Find the conditions behind the average.
A strong average can hide weak performance on particular accents, tones, languages, or acoustic conditions. Break down the results to identify what needs attention.
Qwen3-TTS-12Hz-1.7B · Tone instruction
Pinpoint where your model fails.
Isolate the specific combinations to see where your model needs improvement.
Behind the benchmarks
Explore the research.
The questions, methods, and findings behind a more complete view of voice AI.
Real World VoiceEQ
Introducing Real World VoiceEQ
How the benchmark compares leading voice models across four modalities and more than a million human judgments.
Read the analysisMethodology
Why measuring voice is hard, and why standard approaches aren’t working
What a single preference score leaves out, and what voice evaluation requires instead.
Read the analysisSpeech recognition
Measuring benchmark optimization in speech recognition
How held-out conditions separate genuine capability from familiarity with a public test set.
Read the analysisVoice controllability
Voice controllability benchmark analysis
Whether models can create a requested voice, follow delivery instructions, and hold a defined role.
Read the analysisVoice replication
Voice replication benchmark analysis
Separating speaker-identity fidelity from naturalness and acoustic quality across reference voices.
Read the analysis
Your aggregate score is hiding something. Find out what.
Get in touch with Hume’s research team.