All blog posts

Introducing the Hume Voice Replication Leaderboard

SR
KL
Alice Baird
Sharath Rao, Kimberly Lo, and Alice Baird
··article
Share
Blog Image Replicated Waveforms

Voice replication allows a model to reproduce a specific person’s voice from a short reference recording. For many applications, clear, accurate, and natural speech is not enough: the output must still sound recognizably like the person it is meant to replicate.

Often called voice cloning, the capability can help a creator translate their work without sounding like someone else, let a studio produce an actor-approved pickup after recording has ended, preserve the voice of someone losing the ability to speak, or keep an authorized spokesperson’s voice consistent across customer interactions.

These use cases depend on three distinct qualities: sounding like a particular person, sounding natural, and delivering clear audio.

A new VoiceEQ benchmark for voice replication

Today we are publishing the Voice Replication Leaderboard,a new component of Hume’s Real World VoiceEQ Benchmark. It measures how well leading text-to-speech models reproduce a specific voice across varied real-world conditions.

Like Real-World VoiceEQ, the leaderboard combines human judgments with objective measures rather than reducing performance to a single score.

How we evaluated voice replication

We evaluated eleven leading models, asking each to reproduce the same 25 reference voices. To better reflect real-world use, the reference set spans three categories designed to stress different aspects of replication:

  • Five standard references drawn from archival recordings of public figures and conversational US speech, not specifically selected for accent or emotional variation;
  • Five expressive references spanning anger, fear, happy, sad, and calm
reference_angry_female_expresso
0:00
0:00
reference_fearful_female_expresso
0:00
0:00
reference_happy_male_expresso
0:00
0:00
reference_sad_female_expresso
0:00
0:00
reference_calm_male_expresso
0:00
0:00
  • 15 references selected for accent variation, split into 6 native English varieties and 9 non-native accents, drawn from a mix of read and conversational speech.
reference_french_female
0:00
0:00
reference_english_newcastle_female
0:00
0:00
reference_conv_nigerian_male
0:00
0:00
reference_conv_indian_female
0:00
0:00

Every tested model was asked to replicate these reference voices with the same seven prompts, written to avoid prescribing a particular emotion or delivery style. Each generated clip was evaluated across four dimensions:

Same speaker: A 1–5 human rating of whether the generated clip sounded like the reference speaker.

Quality: A 1–5 human rating of whether the audio was clean and free of artifacts, independent of speaker identity.

Naturalness: A 1–5 human rating of how human the delivery sounded, including pacing and prosody, independent of speaker identity.

Objective similarity: The cosine similarity between the TitaNet speaker embeddings of the reference and generated clips, scored from 0 to 1.

For the three human-rated dimensions, each clip was independently assessed by three paid raters via the Hume Human Feedback API. Raters were blind to the model that produced it and heard the reference recording first, followed by the generated clip. When we report a model's score across all references, we weight the three reference categories equally, one third each, rather than averaging over samples.

Results

Overall, no model was strongest across every dimension or reference type. Performance depended both on what was measured and on the voice being replicated. The results reinforce a core principle of VoiceEQ: voice systems should be evaluated as a profile of capabilities, not reduced to a single score.

Voice Replication Leaderboard
Human raters were asked to score 1–5 of whether a clip generated by a TTS model sounded like the reference speaker provided. Best value per column in accent color. Hover a score for the full row.
#ModelSame speaker ▼QualityNaturalnessObjective sim.
1openbmb/VoxCPM24.214.273.960.84
2meituan-longcat/LongCat-AudioDiT-3.5B4.064.124.020.88
3fishaudio/s2-pro4.034.383.920.74
4bosonai/higgs-audio-v33.834.383.900.74
5cartesia/sonic-3.53.704.334.180.72
6elevenlabs/eleven-multilingual-v23.684.423.900.69
7qwen/Qwen3-TTS-12Hz-1.7B3.654.333.770.71
8cartesia/sonic-3.6-beta3.634.484.360.74
9inworld/inworld-tts-23.624.614.110.64
10microsoft/VibeVoice-1.5B3.463.853.830.70
11elevenlabs/eleven-v32.914.183.900.55
Ranked by same speaker. No model leads every column: identity, quality, naturalness and objective similarity are led by four different models.hume.ai/rw-voice-eq

Sounding natural is not the same as sounding like the right person

A voice can sound convincingly human without sounding like the person it is intended to replicate. Cartesia sonic-3.6-beta received the highest naturalness score (4.36), but ranked eighth on same-speaker similarity (3.63). OpenBMB VoxCPM2 ranked first on same-speaker similarity (4.21), but placed fifth on naturalness (3.96).

This is not evidence of an unavoidable trade-off; for example some models perform well on both. But it shows why the dimensions must be measured separately.

Performance changes with the voice being replicated

Replicating speech from expressive reference voices created the wide gap in performance: Higgs Audio v3 scored 3.38 on emotional references, while VoxCPM2 scored 4.38 on the same category. However no model performed equally well across standard, expressive, and accented references. Notably, Higgs Audio v3 tied with VoxCPM2 for the lead on accented voices, but also had the widest spread across the three reference categories.

Same-speaker score by reference type
Human rating out of 5. The 25 reference voices split into 5 standard, 5 emotional and 15 accented, and the accented set splits again into 6 native English varieties and 9 non-native accents. Darker is higher. Spread is best category minus worst. Best value per column in accent.
#ModelStandardEmotionalAccentNativeNon-nativeSpread
1openbmb/VoxCPM23.984.384.274.174.340.40
2meituan-longcat/LongCat-AudioDiT-3.5B4.063.944.194.094.250.25
3fishaudio/s2-pro3.794.134.174.274.110.38
4bosonai/higgs-audio-v33.823.384.284.354.230.90
5cartesia/sonic-3.53.913.453.743.963.600.46
6elevenlabs/eleven-multilingual-v23.603.963.483.693.340.48
7qwen/Qwen3-TTS-12Hz-1.7B3.663.743.563.663.500.18
8cartesia/sonic-3.6-beta3.673.813.413.593.300.40
9inworld/inworld-tts-23.503.543.833.933.760.33
10microsoft/VibeVoice-1.5B3.413.383.603.403.730.22
11elevenlabs/eleven-v32.822.903.013.172.890.19
The two models at the top score higher on non-native accents than native. Every other model but one does the reverse. Higgs Audio v3 leads accents at 4.28 and drops to 3.38 on emotional references, the widest spread on the board.hume.ai/rw-voice-eq

Clean audio is increasingly becoming table stakes among the models tested

Quality scores were high and tightly grouped, ranging from 3.85 to 4.61, with an average of approximately 4.30. Same-speaker ratings varied more widely, from 2.91 to 4.21. For example, Inworld TTS-2 received the highest quality score (4.61), but ranked ninth on same-speaker similarity (3.62).

Objective and human-rated similarity measures are complementary, not interchangeable

The objective embedding score broadly supported what human listeners heard but the rankings diverged in meaningful cases. VoxCPM2 and LongCat took the top two places under both approaches, in opposite order, and Fish Audio s2-pro and Higgs Audio v3 held third and fourth under both. ElevenLabs Eleven v3 ranked last under both.

Cartesia sonic-3.6-beta ranked eighth by human same-speaker rating and fifth by objective similarity. Speaker embeddings provide a useful, scalable signal, but they do not capture everything people hear when deciding whether two clips sound like the same person. For evaluation and model improvement, objective measures and human judgments are most useful together.

What this means for teams building and choosing voice models

For product teams: test the voice you actually need - given there is no single model that is best across speaker identity, naturalness, quality, and every reference type. The right choice depends on the application. A model suited to neutral narration may not be the best fit for emotionally varied customer service; one that performs well for one speaker may be less reliable for speakers with different accents.

Teams should test models using voices, prompts, languages, emotional conditions, and recording environments that resemble their actual deployment rather than relying on a general score alone.

For model developers: separate the failure modes to drive improvement. Improving voice replication requires measuring identity, naturalness, quality, and objective similarity independently. Researchers should break results down by reference type and, where possible, by individual emotion and accent. This helps determine whether a weakness reflects emotional speech broadly, a particular expressive condition, or a specific type of reference recording.


For example, our leaderboard suggests that emotionally expressive speech remains a challenge. To diagnose where the failure occurs, researchers can test two conditions separately (i) replicating a voice from an emotionally expressive reference while keeping the target text neutral, and (ii) asking the model to deliver emotionally charged text from a neutral reference. Comparing same-speaker and naturalness scores across these conditions can help distinguish identity extraction from expressive generation. If speaker identity remains stable but naturalness falls when emotional delivery is required, the weakness is likely in how the model renders expressive speech. If identity falls when the reference itself is emotional, even with neutral target text, the model may be struggling to recover a consistent representation of the speaker from expressive audio.

Explore the leaderboard

The Voice Replication Leaderboard is live on Hugging Face, with per-category breakdowns and curated reference and replica audio samples. It extends the Real World VoiceEQ Benchmark’s effort to evaluate voice AI the way people actually experience it.

Public leaderboards provide a shared baseline. Hume also works with frontier labs and product teams to design evaluations around their own speakers, languages, emotional contexts, and production failure modes—combining human ratings and objective measures to identify weaknesses and generate the preference data needed to improve them. If you’re evaluating a model or developing one yourself, get in touch,

Stay in the loop

Get the latest on empathic AI research, product updates, and company news.

Join the community

Connect with other developers, share projects, and get help from the team.

Join our Discord