Now liveExplore Expression APIs
hume.ai logo
Skip to the benchmark ↓

Real World VoiceEQ

Voice AI has made remarkable progress.

Faster responses.Fewer errors.

Something still feels off.

Scroll

TTS leaderboard

Text-to-speech, scored on what it produces

We score how well a model turns text into speech across the dimensions that decide real-world quality: expressive range, identity stability, pronunciation accuracy, and audio cleanliness. What matters is what the model produces across a full performance, not how clean one sentence sounds.

Overall

Frontier score: closeness to the ideal corner (TOPSIS) on two rater-controlled axes, Expressivity and Reliability, each built from the individual rater questions that measure it across all evals. Models without an instruction channel read the same prompts without the direction. 0-1; higher is better.

  1. 01google/gemini-3.8-flashProprietary0.92
  2. 02google/gemini-3.8-flash-liteProprietary0.91
  3. 03google/gemini-2.5-proProprietary0.88
  4. 04google/gemini-2.5-flashProprietary0.86
  5. 05cartesia/sonic-3.6Proprietary0.84
  6. 06xai/grok-ttsProprietary0.79
  7. 07google/gemini-3.1-flashProprietary0.78
  8. 08elevenlabs/eleven_v3_conversationalProprietary0.77
  9. 09cartesia/sonic-3.5Proprietary0.76
  10. 10speechify/simba-3.2Proprietary0.75
  11. 11openai/gpt-4o-mini-ttsProprietary0.74
  12. 12fishaudio/s2-proOpen0.72
  13. 13elevenlabs/eleven_v3Proprietary0.71
  14. 14openai/tts-1Proprietary0.69
  15. 15openai/tts-1-hdProprietary0.69
  16. 16inworld/tts-2-flashProprietary0.68
  17. 17inworld/tts-1-maxProprietary0.65
  18. 18elevenlabs/eleven_multilingual_v2Proprietary0.64
  19. 19deepgram/aura-2Proprietary0.59
  20. 20smallest/lightning_v3.1Proprietary0.58
  21. 21inworld/tts-2Proprietary0.58
  22. 22elevenlabs/eleven_flash_v2_5Proprietary0.57
  23. 23microsoft/VibeVoice-1.5BOpen0.56
  24. 24hexgrad/Kokoro-82MOpen0.54
  25. 25inworld/tts-1Proprietary0.54
  26. 26resemble-ai/chatterboxOpen0.51
  27. 27qwen/Qwen3-TTS-12Hz-1.7BOpen0.49
  28. 28hume/tada-3b-mlOpen0.48
  29. 29sesame/csm-1bOpen0.38
  30. 30coqui/xtts-v2Open0.37
  31. 31bosonai/higgs-audio-v2Open0.35
  32. 32bosonai/higgs-audio-v3-tts-4bOpen0.29
  33. 33indexteam/indextts-2Open0.18
  34. 34nari-labs/diaOpen0.17
  35. 35parler-tts/parler-ttsOpen0.04

Explore the complete results.

View full model rankings, audio samples, methodology, per-dimension results, and additional benchmark analysis.

Private evaluations

Evaluate your system in the conditions that matter.

A public benchmark shows how you compare at a high level. A private evaluation gives you specific data and insights into where your system falls short and what to improve.

Built on the same research and evaluation infrastructure, tailored to your system and use case.

  • Your system

    Model, agent, endpoint, or checkpoint

  • Your criteria

    Tailored to the modality and use case

  • Your conditions

    Scenarios, speakers, and environments relevant to your users

  • Your results, private

    Condition-level analysis, kept to your team

  • Checkpoint comparison

    Internal checkpoints or external systems, side by side

  • Reusable regression

    A held-out suite to rerun on the next release

Request a private evaluation

How we measure

The right judge for every measure.

Different qualities require different ways of measuring them. Hume matches each metric to objective measures, model graders, or trained human raters, combining approaches where needed.

  1. 01

    Objective measures

    • Accuracy
    • Timing
    • Acoustics
  2. 02

    Automated graders

    • Defined criteria
    • At scale
    • Checked against human ratings
  3. 03

    Human raters

    • Naturalness
    • Expression
    • Context
Request a private evaluation

What a private evaluation shows

Get the diagnostic-level detail to guide what you do next

A private evaluation returns capability-level results for your own system, and the conditions underneath each average.

Find the conditions behind the average.

A strong average can hide weak performance on particular accents, tones, languages, or acoustic conditions. Break down the results to identify what needs attention.

Qwen3-TTS-12Hz-1.7B · Tone instruction

Overall tone 63.8%
Formality 47.6%
Warmth 54.0%
Sarcasm 55.6%
Intensity 75.0%
Authority 84.9%

Pinpoint where your model fails.

Isolate the specific combinations to see where your model needs improvement.

Naturalness
German(passes)
Male(passes)Female(fails)
Anger(fails)
Joy(fails)
Sadness(passes)
Fear(passes)
Surprise(fails)
Disgust(passes)
Warmth(fails)
Sarcasm(passes)
Intensity(passes)
Request a private evaluation

Behind the benchmarks

Explore the research.

The questions, methods, and findings behind a more complete view of voice AI.

  • Real World VoiceEQ

    Introducing Real World VoiceEQ

    How the benchmark compares leading voice models across four modalities and more than a million human judgments.

    Read the analysis
  • Methodology

    Why measuring voice is hard, and why standard approaches aren’t working

    What a single preference score leaves out, and what voice evaluation requires instead.

    Read the analysis
  • Speech recognition

    Measuring benchmark optimization in speech recognition

    How held-out conditions separate genuine capability from familiarity with a public test set.

    Read the analysis
  • Voice controllability

    Voice controllability benchmark analysis

    Whether models can create a requested voice, follow delivery instructions, and hold a defined role.

    Read the analysis
  • Voice replication

    Voice replication benchmark analysis

    Separating speaker-identity fidelity from naturalness and acoustic quality across reference voices.

    Read the analysis

Your aggregate score is hiding something. Find out what.

Get in touch with Hume’s research team.