
Voice replication allows a model to reproduce a specific person’s voice from a short reference recording. For many applications, clear, accurate, and natural speech is not enough: the output must still sound recognizably like the person it is meant to replicate.
Often called voice cloning, the capability can help a creator translate their work without sounding like someone else, let a studio produce an actor-approved pickup after recording has ended, preserve the voice of someone losing the ability to speak, or keep an authorized spokesperson’s voice consistent across customer interactions.
These use cases depend on three distinct qualities: sounding like a particular person, sounding natural, and delivering clear audio.
A new VoiceEQ benchmark for voice replication
Today we are publishing the Voice Replication Leaderboard,a new component of Hume’s Real World VoiceEQ Benchmark. It measures how well leading text-to-speech models reproduce a specific voice across varied real-world conditions.
Like Real-World VoiceEQ, the leaderboard combines human judgments with objective measures rather than reducing performance to a single score.
How we evaluated voice replication
We evaluated eleven leading models, asking each to reproduce the same 25 reference voices. To better reflect real-world use, the reference set spans three categories designed to stress different aspects of replication:
- Five standard references drawn from archival recordings of public figures and conversational US speech, not specifically selected for accent or emotional variation;
- Five expressive references spanning anger, fear, happy, sad, and calm
- 15 references selected for accent variation, split into 6 native English varieties and 9 non-native accents, drawn from a mix of read and conversational speech.
Every tested model was asked to replicate these reference voices with the same seven prompts, written to avoid prescribing a particular emotion or delivery style. Each generated clip was evaluated across four dimensions:
Same speaker: A 1–5 human rating of whether the generated clip sounded like the reference speaker.
Quality: A 1–5 human rating of whether the audio was clean and free of artifacts, independent of speaker identity.
Naturalness: A 1–5 human rating of how human the delivery sounded, including pacing and prosody, independent of speaker identity.
Objective similarity: The cosine similarity between the TitaNet speaker embeddings of the reference and generated clips, scored from 0 to 1.
For the three human-rated dimensions, each clip was independently assessed by three paid raters via the Hume Human Feedback API. Raters were blind to the model that produced it and heard the reference recording first, followed by the generated clip. When we report a model's score across all references, we weight the three reference categories equally, one third each, rather than averaging over samples.
Results
Overall, no model was strongest across every dimension or reference type. Performance depended both on what was measured and on the voice being replicated. The results reinforce a core principle of VoiceEQ: voice systems should be evaluated as a profile of capabilities, not reduced to a single score.
| # | Model | Same speaker ▼ | Quality | Naturalness | Objective sim. |
|---|---|---|---|---|---|
| 1 | openbmb/VoxCPM2 | 4.21 | 4.27 | 3.96 | 0.84 |
| 2 | meituan-longcat/LongCat-AudioDiT-3.5B | 4.06 | 4.12 | 4.02 | 0.88 |
| 3 | fishaudio/s2-pro | 4.03 | 4.38 | 3.92 | 0.74 |
| 4 | bosonai/higgs-audio-v3 | 3.83 | 4.38 | 3.90 | 0.74 |
| 5 | cartesia/sonic-3.5 | 3.70 | 4.33 | 4.18 | 0.72 |
| 6 | elevenlabs/eleven-multilingual-v2 | 3.68 | 4.42 | 3.90 | 0.69 |
| 7 | qwen/Qwen3-TTS-12Hz-1.7B | 3.65 | 4.33 | 3.77 | 0.71 |
| 8 | cartesia/sonic-3.6-beta | 3.63 | 4.48 | 4.36 | 0.74 |
| 9 | inworld/inworld-tts-2 | 3.62 | 4.61 | 4.11 | 0.64 |
| 10 | microsoft/VibeVoice-1.5B | 3.46 | 3.85 | 3.83 | 0.70 |
| 11 | elevenlabs/eleven-v3 | 2.91 | 4.18 | 3.90 | 0.55 |
Sounding natural is not the same as sounding like the right person
A voice can sound convincingly human without sounding like the person it is intended to replicate. Cartesia sonic-3.6-beta received the highest naturalness score (4.36), but ranked eighth on same-speaker similarity (3.63). OpenBMB VoxCPM2 ranked first on same-speaker similarity (4.21), but placed fifth on naturalness (3.96).
This is not evidence of an unavoidable trade-off; for example some models perform well on both. But it shows why the dimensions must be measured separately.
Performance changes with the voice being replicated
Replicating speech from expressive reference voices created the wide gap in performance: Higgs Audio v3 scored 3.38 on emotional references, while VoxCPM2 scored 4.38 on the same category. However no model performed equally well across standard, expressive, and accented references. Notably, Higgs Audio v3 tied with VoxCPM2 for the lead on accented voices, but also had the widest spread across the three reference categories.
| # | Model | Standard | Emotional | Accent | Native | Non-native | Spread |
|---|---|---|---|---|---|---|---|
| 1 | openbmb/VoxCPM2 | 3.98 | 4.38 | 4.27 | 4.17 | 4.34 | 0.40 |
| 2 | meituan-longcat/LongCat-AudioDiT-3.5B | 4.06 | 3.94 | 4.19 | 4.09 | 4.25 | 0.25 |
| 3 | fishaudio/s2-pro | 3.79 | 4.13 | 4.17 | 4.27 | 4.11 | 0.38 |
| 4 | bosonai/higgs-audio-v3 | 3.82 | 3.38 | 4.28 | 4.35 | 4.23 | 0.90 |
| 5 | cartesia/sonic-3.5 | 3.91 | 3.45 | 3.74 | 3.96 | 3.60 | 0.46 |
| 6 | elevenlabs/eleven-multilingual-v2 | 3.60 | 3.96 | 3.48 | 3.69 | 3.34 | 0.48 |
| 7 | qwen/Qwen3-TTS-12Hz-1.7B | 3.66 | 3.74 | 3.56 | 3.66 | 3.50 | 0.18 |
| 8 | cartesia/sonic-3.6-beta | 3.67 | 3.81 | 3.41 | 3.59 | 3.30 | 0.40 |
| 9 | inworld/inworld-tts-2 | 3.50 | 3.54 | 3.83 | 3.93 | 3.76 | 0.33 |
| 10 | microsoft/VibeVoice-1.5B | 3.41 | 3.38 | 3.60 | 3.40 | 3.73 | 0.22 |
| 11 | elevenlabs/eleven-v3 | 2.82 | 2.90 | 3.01 | 3.17 | 2.89 | 0.19 |
Clean audio is increasingly becoming table stakes among the models tested
Quality scores were high and tightly grouped, ranging from 3.85 to 4.61, with an average of approximately 4.30. Same-speaker ratings varied more widely, from 2.91 to 4.21. For example, Inworld TTS-2 received the highest quality score (4.61), but ranked ninth on same-speaker similarity (3.62).
Objective and human-rated similarity measures are complementary, not interchangeable
The objective embedding score broadly supported what human listeners heard but the rankings diverged in meaningful cases. VoxCPM2 and LongCat took the top two places under both approaches, in opposite order, and Fish Audio s2-pro and Higgs Audio v3 held third and fourth under both. ElevenLabs Eleven v3 ranked last under both.
Cartesia sonic-3.6-beta ranked eighth by human same-speaker rating and fifth by objective similarity. Speaker embeddings provide a useful, scalable signal, but they do not capture everything people hear when deciding whether two clips sound like the same person. For evaluation and model improvement, objective measures and human judgments are most useful together.
What this means for teams building and choosing voice models
For product teams: test the voice you actually need - given there is no single model that is best across speaker identity, naturalness, quality, and every reference type. The right choice depends on the application. A model suited to neutral narration may not be the best fit for emotionally varied customer service; one that performs well for one speaker may be less reliable for speakers with different accents.
Teams should test models using voices, prompts, languages, emotional conditions, and recording environments that resemble their actual deployment rather than relying on a general score alone.
For model developers: separate the failure modes to drive improvement. Improving voice replication requires measuring identity, naturalness, quality, and objective similarity independently. Researchers should break results down by reference type and, where possible, by individual emotion and accent. This helps determine whether a weakness reflects emotional speech broadly, a particular expressive condition, or a specific type of reference recording.
For example, our leaderboard suggests that emotionally expressive speech remains a challenge. To diagnose where the failure occurs, researchers can test two conditions separately (i) replicating a voice from an emotionally expressive reference while keeping the target text neutral, and (ii) asking the model to deliver emotionally charged text from a neutral reference. Comparing same-speaker and naturalness scores across these conditions can help distinguish identity extraction from expressive generation. If speaker identity remains stable but naturalness falls when emotional delivery is required, the weakness is likely in how the model renders expressive speech. If identity falls when the reference itself is emotional, even with neutral target text, the model may be struggling to recover a consistent representation of the speaker from expressive audio.
Explore the leaderboard
The Voice Replication Leaderboard is live on Hugging Face, with per-category breakdowns and curated reference and replica audio samples. It extends the Real World VoiceEQ Benchmark’s effort to evaluate voice AI the way people actually experience it.
Public leaderboards provide a shared baseline. Hume also works with frontier labs and product teams to design evaluations around their own speakers, languages, emotional contexts, and production failure modes—combining human ratings and objective measures to identify weaknesses and generate the preference data needed to improve them. If you’re evaluating a model or developing one yourself, get in touch,



