
Text-to-speech systems were originally built to read text aloud in a single voice. As voice AI expands into audiobooks, game dialogue, and advertising, these systems are being asked to do more: generate conversations between multiple speakers.
That takes more than generating two distinct voices and stitching their lines together. The speakers need to sound like they are responding to one another, with natural timing, changes in tone, and smooth handoffs. Each voice must remain distinct and consistent while contributing to a believable conversation.
We independently evaluated Google’s newest models for our public Real-World VoiceEQ benchmark. Separately, we developed several bespoke evals including a rubric for the multi-speaker dialogue, a functionality now becoming more available in commercial TTS models. With those rubrics in place, we could assess progress across model generations and identify opportunities for further improvement.
How we evaluated multi-speaker dialogue
Traditional metrics such as word error rate (WER) focus on whether a system gets the words right. With Real-World VoiceEQ, we developed more holistic measures that assess both the objective and perceptual qualities of voice and expression.
Building on that work, we developed rubrics to assess speaker identity, conversational flow, and emotional delivery in multi-speaker dialogue. We tested these qualities through four scenarios designed to expose specific ways a conversation can break down:
Similar-sounding speakers. Conversations between people with similar pitch and timbre test the model’s ability to keep their voices distinct throughout.
Frequent interruptions. Short, rapid turns and speakers interrupting each other test the model’s ability to handle transitions naturally, without audible artifacts or awkward pauses.
Contrasting emotions. Dialogue where one speaker is agitated while the other remains calm tests the model’s ability to preserve each speaker’s emotional delivery without one bleeding into the other.
Longer dialogue. Conversations spanning 5, 10, and 15 turns test the model’s ability to maintain each speaker’s voice identity as the dialogue gets longer.
We developed 48 two-speaker scripts based on these scenarios. Each model generated audio for every script using male-female, female-female, and male-male speaker combinations.
Using these rubrics, three independent raters evaluated every clip through the Hume Human Feedback API, without knowing which model produced it. They scored each clip on a 1-to-5 scale:
- Speaker separation. Could you tell the two speakers apart throughout, without effort?
- Voice stability. Does each speaker's voice stay the same from start to finish?
- Seamlessness. Does this sound like one continuous recording, or like separate clips joined together?
- Conversation dynamics. Ignoring what is said, does the back-and-forth sound like two people interacting?
- Emotional containment. Did each speaker keep the emotional state their own lines call for, or did one speaker's mood spill into the other's?
Using private evaluations to guide model development
With the rubrics established, we analyse the human ratings and compare model generations using consistent criteria to assess progress and identify areas for further improvement.
We evaluated five Gemini TTS models across three generations. The table lists them by model family and version; Overall is the mean of the five dimensions.
| Model | Speaker separation | Voice stability | Seamlessness | Conversation dynamics | Emotional containment | Overall |
|---|---|---|---|---|---|---|
| Gemini 2.5 Flash TTS | 3.24 | 4.18 | 4.10 | 3.61 | 3.88 | 3.80 |
| Gemini 2.5 Pro TTS | 3.90 | 4.14 | 4.24 | 4.06 | 4.26 | 4.12 |
| Gemini 3.1 Flash TTS | 3.08 | 3.99 | 3.91 | 3.43 | 3.60 | 3.60 |
| Gemini 3.8 Flash TTS | 4.17 | 4.04 | 3.97 | 4.08 | 4.29 | 4.11 |
| Gemini 3.8 Flash-Lite TTS | 3.88 | 4.14 | 4.21 | 3.96 | 4.31 | 4.10 |
Emotional containment is a strength and is the highest-scoring dimension for Gemini 2.5 Pro, 3.8 Flash and 3.8 Flash-Lite, at 4.26 to 4.31. When one speaker is agitated and the other calm, these models keep each speaker’s delivery their own.
Speaker separation is the clearest area for improvement. It is the lowest-scoring dimension for four of the five models, averaging 3.65 against roughly 4.1 for voice stability, seamlessness and emotional containment. Gemini 3.8 Flash is the exception, at 4.17.
Speaker separation: same-gender pairs are the most challenging
| Model | Male-female | Female-female | Male-male | Separation overall |
|---|---|---|---|---|
| Gemini 2.5 Flash TTS | 4.88 | 3.08 | 1.75 | 3.24 |
| Gemini 2.5 Pro TTS | 4.77 | 4.17 | 2.77 | 3.90 |
| Gemini 3.1 Flash TTS | 3.79 | 2.31 | 3.12 | 3.08 |
| Gemini 3.8 Flash TTS | 4.92 | 4.17 | 3.42 | 4.17 |
| Gemini 3.8 Flash-Lite TTS | 4.88 | 3.46 | 3.31 | 3.88 |
Across Gemini models, mixed-gender pairs average 4.65 on speaker separation, and four of the five score 4.77 or higher.
Same-gender pairs tell a different story: female-female pairs average 3.44 and male-male pairs 2.87. Male-male is the weakest pairing for every Gemini model except 3.1 Flash, which is weakest on female-female (2.31). Gemini 2.5 Flash TTS scores 4.88 on mixed-gender pairs and 1.75 on male-male pairs, a gap of more than three points on a five-point scale.
Gemini 2.5 Flash also scores 0.2 higher overall than 3.1 Flash while scoring more than a point lower on male-male pairs, the kind of difference an overall column hides. On male-male speaker separation, Gemini 3.8 Flash scores 3.42, compared with 1.75 for Gemini 2.5 Flash.
These results show how targeted evaluations can reveal both progress and priorities for the next iteration. Gemini 3.8 Flash improved speaker separation on male-male pairs compared with earlier Flash models, while the gap with mixed-gender pairs highlights an opportunity for further improvement. Testing these conditions explicitly gives researchers a specific capability to investigate and a consistent way to assess whether subsequent changes improve it.
Tracking progress and trade-offs across generations
Gemini 2.5 Pro, 3.8 Flash and 3.8 Flash-Lite score within 0.02 of each other overall, but get there differently. Gemini 3.8 Flash is stronger on speaker separation (4.17, against 3.90 and 3.88). Gemini 2.5 Pro and 3.8 Flash-Lite are stronger on seamlessness (4.24 and 4.21, against 3.97). These separate measures reveal development priorities that would be difficult to identify from the overall scores alone.
Within the Flash line, Gemini 3.8 Flash improves on every dimension over 3.1 Flash, which scored below 2.5 Flash overall (3.60 against 3.80). Compared with 2.5 Flash, 3.8 Flash gains on speaker separation (+0.93), conversation dynamics (+0.47) and emotional containment (+0.41), and slips slightly on seamlessness (−0.13) and voice stability (−0.14). Tracking these dimensions separately helps researchers identify gains while flagging possible trade-offs in seamlessness and voice stability for further testing.
Putting model performance in context
Comparisons with other systems help researchers distinguish challenges shared across models from weaknesses that vary by model. We applied the same scripts, speaker pairings, and rating questions to other multi-speaker TTS systems to put the Gemini findings in that context.
Comparisons with other models showed similar same-gender performance gaps, suggesting this challenge is not isolated to Gemini. All three systems score 4.38 or higher on mixed-gender pairs and drop on at least one same-gender pairing. ElevenLabs v3 shows the opposite pattern to most Gemini models: it scores higher on male-male pairs (4.33) than on female-female pairs (3.02).
See model comparison performance
| Model | Speaker separation | Voice stability | Seamlessness | Conversation dynamics | Emotional containment | Overall |
|---|---|---|---|---|---|---|
| Gemini 2.5 Pro TTS | 3.90 | 4.14 | 4.24 | 4.06 | 4.26 | 4.12 |
| Gemini 3.8 Flash TTS | 4.17 | 4.04 | 3.97 | 4.08 | 4.29 | 4.11 |
| Gemini 3.8 Flash-Lite TTS | 3.88 | 4.14 | 4.21 | 3.96 | 4.31 | 4.10 |
| ElevenLabs v3 | 4.06 | 4.31 | 3.51 | 3.47 | 4.15 | 3.90 |
| Gemini 2.5 Flash TTS | 3.24 | 4.18 | 4.10 | 3.61 | 3.88 | 3.80 |
| Gemini 3.1 Flash TTS | 3.08 | 3.99 | 3.91 | 3.43 | 3.60 | 3.60 |
| Higgs Audio v2 (3B) | 3.40 | 4.03 | 3.38 | 3.10 | 3.47 | 3.48 |
| VibeVoice 1.5B † | 3.35 | 3.38 | 3.39 | 3.19 | 3.62 | 3.39 |
See model comparison performance for gender pairs
| Model | Male-female | Female-female | Male-male | Separation overall |
|---|---|---|---|---|
| Gemini 3.8 Flash TTS | 4.92 | 4.17 | 3.42 | 4.17 |
| ElevenLabs v3 | 4.83 | 3.02 | 4.33 | 4.06 |
| Gemini 2.5 Pro TTS | 4.77 | 4.17 | 2.77 | 3.90 |
| Gemini 3.8 Flash-Lite TTS | 4.88 | 3.46 | 3.31 | 3.88 |
| Higgs Audio v2 (3B) | 4.90 | 2.29 | 3.02 | 3.40 |
| VibeVoice 1.5B | 4.38 | 3.06 | 2.62 | 3.35 |
| Gemini 2.5 Flash TTS | 4.88 | 3.08 | 1.75 | 3.24 |
| Gemini 3.1 Flash TTS | 3.79 | 2.31 | 3.12 | 3.08 |
ElevenLabs v3 scores 4.06 on speaker separation and 3.47 on conversation dynamics, the largest gap between the two of any system we evaluated. Keeping two voices apart and making an exchange sound like two people talking are different capabilities, and a system can be strong at one while trailing on the other.
Why new capabilities need new evaluations
The work began with defining what good multi-speaker dialogue looks like and developing rubrics to measure it. Applying those rubrics across model generations made progress and remaining weaknesses visible, including speaker-separation problems that overall scores could conceal. The resulting evaluation gives researchers a repeatable way to identify priorities for the next iteration and assess whether subsequent changes improve performance.
We help research teams develop bespoke rubrics and provide the infrastructure to run private evaluations. If you are developing a new capability, get in touch to discuss how to define quality and measure progress.



![Blog Image [whisper] Tag Applied](/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2Fxqnc2for%2Fproduction%2F06414e1fae798399961068290af810114be8148e-1200x800.jpg%3Fw%3D1200%26h%3D800&w=3840&q=75)