
Most text to speech benchmarks score a single sentence read aloud. Almost nothing anyone ships is a single sentence. The differences that matter now are harder to hear: whether a voice holds across ten minutes of output, whether it performs the emotion a line calls for or only announces it, whether it survives an unfamiliar name in a language it was not built around.
In July, we released Real-World VoiceEQ, a first-of-its-kind benchmark designed to evaluate the human dimensions of voice – from tone, expression, to naturalness and speaker identity across.
As use cases for voice models become more widespread and complex, we introduced additional evaluation benchmarks covering TTS models, including voice controllability, voice replication and more recently a study into evaluating multi-speaker dialogue. As models improve and evolve, so do the tools that measure what “good” sounds like – in this case public benchmarks and the private evaluations that researchers use to measure their own progress.
And as new models are released by leading labs, Hume researchers conduct independent assessments of the models either using publicly available end-points or via a preview made available by the model developers. We run each model through a consistent set of evaluations, human ratings studies, and comparative analysis. All runs via our Kairos evaluations platform, are done by vetted evaluators who judge results without knowing which models are used – you can learn more about our study, the dimensions measured, and methodology here.
How Google’s latest TTS voice models, Gemini 3.8 Flash and Flash-Lite, measures up
In this release we’ve also introduced a new measure called “Frontier Metric” to the Real World Voice-EQ benchmark. While each of our metrics cover a single dimension (e.g. expressivity, role fit, voice identity etc), frontier metrics are the culmination of two key axis: expressivity and reliability. That’s because through our evaluation we see a clear theme on model performance; more expressive models tend to make that gain at the expense of reliability and we believe being able to make improvements across both is going to define this next chapter of voice AI performance. To that end, we believe that a reasonable proxy for what constitutes frontier voice models is the ability to excel on both at the same time.

On this measure, the most reliable models, grok and gpt-4o-mini, are well down on expressivity, and the most expressive incumbents, Cartesia 3.6 and the previous Gemini models, give up more reliability. Showing that the frontier is advancing, Gemini 3.8 Flash is at the top of the expressivity axis while also sitting in the upper cohort on reliability performance, taking first place on this measure.
- Overall TTS performance: higher scores, with the biggest gains in long-form stability. Flash and Flash-Lite improve on Gemini 3.1 Flash’s overall VoiceEQ score, taking first and second place among 34 models. Extended long-form stability rises from 1.22 to 2.89 for Flash and 3.03 for Flash-Lite, with smaller gains in regular and multilingual long-form performance.
- Voice controllability: a broad strength across multiple dimensions. Flash ranks first overall in Voice Design, with leading scores for accent and gender and strong results for texture and prosody. Both models also rank among the leaders at following style tags while keeping speech natural—showing strength in controlling both a voice’s characteristics and its delivery.
- Voice replication: Flash-Lite performs around the field average. Reproducing a specific person’s voice is less differentiated: Flash-Lite ranks seventh of 13 models, with a same-speaker score almost identical to the average.
- Precision of voice customization: room to improve. Despite Flash’s strong overall controllability, some attributes remain harder to get right. Matching a young-sounding voice and controlling volume are clearer weaknesses relative to leading competitors.
Beyond the benchmark: how leading AI labs work with Hume to use evaluation as part of model improvement
As the frontier advances and new capabilities are born, so too, does the tooling to evaluate and improve. A public benchmark is a useful way for the industry, potential users of models wanting to build applications off them, to get a high level overview of performance across the categories that matter most, and private evaluations is where that data becomes intelligence for decision making, whether that is what model is used, or if you’re a researcher, the insights on where exactly to improve the model.
That’s because as comprehensive as our benchmark is, (i) evaluations need to be tailored to use case whether that is domain, modality, or interaction specific and (ii) models are getting so good that failure modes aren’t typically entire failures of dimensions but in specific areas and instances. For example, a model might do well on an overall multilingual category but find that they’re doing well with european languages and below average in asian languages. Or that their below average score on German is due to poor performance on a regional dialect of German. Or that an ASR model can generally do well with background noise if that noise is music, but fails when it’s background chatter.
Private benchmarks bridge that divide and become operational measures to know where exactly to improve and that a training run actually made the model better in the parts where they are trying to improve.
When research labs partner with Hume, we take care to ensure that the prompts and reference material of our benchmark are held out and not shared with the team. This is important in maintaining integrity of the benchmark so it cannot be trained against, and no score can be improved by memorising the test (as we’ve seen in our research on the optimisation of public benchmarks).
Hume’s Kairos platform is what powers private evaluation. Kairos is an end to end solution that turns around large scale evaluations in record time, covering evaluation design, simulated conversations, user testing, expressivity analysis, audio processing and ratings pipeline, with audio measurement, model to market benchmarking, multilingual analysis and regressions, and the gold standard, human ratings.
Get in touch if you’re a researcher wanting to do private evaluation.


![Blog Image [whisper] Tag Applied](/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2Fxqnc2for%2Fproduction%2F06414e1fae798399961068290af810114be8148e-1200x800.jpg%3Fw%3D1200%26h%3D800&w=3840&q=75)
