All blog posts

Evaluating Google’s multi-speaker TTS: A case study in why private evaluations matter

SR
KL
Alice Baird
Sharath Rao, Kimberly Lo, and Alice Baird
··article
Share
Evaluating Google’s multi-speaker TTS: A case study in why private evaluations matter

Text-to-speech systems were originally built to read text aloud in a single voice. As voice AI expands into audiobooks, game dialogue, and advertising, these systems are being asked to do more: generate conversations between multiple speakers.

That takes more than generating two distinct voices and stitching their lines together. The speakers need to sound like they are responding to one another, with natural timing, changes in tone, and smooth handoffs. Each voice must remain distinct and consistent while contributing to a believable conversation.

We independently evaluated Google’s newest models for our public Real-World VoiceEQ benchmark. Separately, we developed several bespoke evals including a rubric for the multi-speaker dialogue, a functionality now becoming more available in commercial TTS models. With those rubrics in place, we could assess progress across model generations and identify opportunities for further improvement.

How we evaluated multi-speaker dialogue

Traditional metrics such as word error rate (WER) focus on whether a system gets the words right. With Real-World VoiceEQ, we developed more holistic measures that assess both the objective and perceptual qualities of voice and expression.

Building on that work, we developed rubrics to assess speaker identity, conversational flow, and emotional delivery in multi-speaker dialogue. We tested these qualities through four scenarios designed to expose specific ways a conversation can break down:

Similar-sounding speakers. Conversations between people with similar pitch and timbre test the model’s ability to keep their voices distinct throughout.

Similar Speakers
0:00
0:00

Frequent interruptions. Short, rapid turns and speakers interrupting each other test the model’s ability to handle transitions naturally, without audible artifacts or awkward pauses.

Frequent Interruptions
0:00
0:00

Contrasting emotions. Dialogue where one speaker is agitated while the other remains calm tests the model’s ability to preserve each speaker’s emotional delivery without one bleeding into the other.

Contrasting Emotions
0:00
0:00

Longer dialogue. Conversations spanning 5, 10, and 15 turns test the model’s ability to maintain each speaker’s voice identity as the dialogue gets longer.

Longer Dialogue
0:00
0:00

We developed 48 two-speaker scripts based on these scenarios. Each model generated audio for every script using male-female, female-female, and male-male speaker combinations.

Using these rubrics, three independent raters evaluated every clip through the Hume Human Feedback API, without knowing which model produced it. They scored each clip on a 1-to-5 scale:

  • Speaker separation. Could you tell the two speakers apart throughout, without effort?
  • Voice stability. Does each speaker's voice stay the same from start to finish?
  • Seamlessness. Does this sound like one continuous recording, or like separate clips joined together?
  • Conversation dynamics. Ignoring what is said, does the back-and-forth sound like two people interacting?
  • Emotional containment. Did each speaker keep the emotional state their own lines call for, or did one speaker's mood spill into the other's?

Using private evaluations to guide model development

With the rubrics established, we analyse the human ratings and compare model generations using consistent criteria to assess progress and identify areas for further improvement.

We evaluated five Gemini TTS models across three generations. The table lists them by model family and version; Overall is the mean of the five dimensions.

ModelSpeaker separationVoice stabilitySeamlessnessConversation dynamicsEmotional containmentOverall
Gemini 2.5 Flash TTS3.244.184.103.613.883.80
Gemini 2.5 Pro TTS3.904.144.244.064.264.12
Gemini 3.1 Flash TTS3.083.993.913.433.603.60
Gemini 3.8 Flash TTS4.174.043.974.084.294.11
Gemini 3.8 Flash-Lite TTS3.884.144.213.964.314.10

Emotional containment is a strength and is the highest-scoring dimension for Gemini 2.5 Pro, 3.8 Flash and 3.8 Flash-Lite, at 4.26 to 4.31. When one speaker is agitated and the other calm, these models keep each speaker’s delivery their own.

Speaker separation is the clearest area for improvement. It is the lowest-scoring dimension for four of the five models, averaging 3.65 against roughly 4.1 for voice stability, seamlessness and emotional containment. Gemini 3.8 Flash is the exception, at 4.17.

Speaker separation: same-gender pairs are the most challenging

ModelMale-femaleFemale-femaleMale-maleSeparation overall
Gemini 2.5 Flash TTS4.883.081.753.24
Gemini 2.5 Pro TTS4.774.172.773.90
Gemini 3.1 Flash TTS3.792.313.123.08
Gemini 3.8 Flash TTS4.924.173.424.17
Gemini 3.8 Flash-Lite TTS4.883.463.313.88

Across Gemini models, mixed-gender pairs average 4.65 on speaker separation, and four of the five score 4.77 or higher.

Separation strong male female
0:00
0:00
Separation strong male male
0:00
0:00
Separation strong female female
0:00
0:00

Same-gender pairs tell a different story: female-female pairs average 3.44 and male-male pairs 2.87. Male-male is the weakest pairing for every Gemini model except 3.1 Flash, which is weakest on female-female (2.31). Gemini 2.5 Flash TTS scores 4.88 on mixed-gender pairs and 1.75 on male-male pairs, a gap of more than three points on a five-point scale.

Gemini 2.5 Flash also scores 0.2 higher overall than 3.1 Flash while scoring more than a point lower on male-male pairs, the kind of difference an overall column hides. On male-male speaker separation, Gemini 3.8 Flash scores 3.42, compared with 1.75 for Gemini 2.5 Flash.

Separation weak male male
0:00
0:00
Separation weak female female
0:00
0:00

These results show how targeted evaluations can reveal both progress and priorities for the next iteration. Gemini 3.8 Flash improved speaker separation on male-male pairs compared with earlier Flash models, while the gap with mixed-gender pairs highlights an opportunity for further improvement. Testing these conditions explicitly gives researchers a specific capability to investigate and a consistent way to assess whether subsequent changes improve it.

Tracking progress and trade-offs across generations

Gemini 2.5 Pro, 3.8 Flash and 3.8 Flash-Lite score within 0.02 of each other overall, but get there differently. Gemini 3.8 Flash is stronger on speaker separation (4.17, against 3.90 and 3.88). Gemini 2.5 Pro and 3.8 Flash-Lite are stronger on seamlessness (4.24 and 4.21, against 3.97). These separate measures reveal development priorities that would be difficult to identify from the overall scores alone.

Within the Flash line, Gemini 3.8 Flash improves on every dimension over 3.1 Flash, which scored below 2.5 Flash overall (3.60 against 3.80). Compared with 2.5 Flash, 3.8 Flash gains on speaker separation (+0.93), conversation dynamics (+0.47) and emotional containment (+0.41), and slips slightly on seamlessness (−0.13) and voice stability (−0.14). Tracking these dimensions separately helps researchers identify gains while flagging possible trade-offs in seamlessness and voice stability for further testing.

Putting model performance in context

Comparisons with other systems help researchers distinguish challenges shared across models from weaknesses that vary by model. We applied the same scripts, speaker pairings, and rating questions to other multi-speaker TTS systems to put the Gemini findings in that context.

Comparisons with other models showed similar same-gender performance gaps, suggesting this challenge is not isolated to Gemini. All three systems score 4.38 or higher on mixed-gender pairs and drop on at least one same-gender pairing. ElevenLabs v3 shows the opposite pattern to most Gemini models: it scores higher on male-male pairs (4.33) than on female-female pairs (3.02).

See model comparison performance
ModelSpeaker separationVoice stabilitySeamlessnessConversation dynamicsEmotional containmentOverall
Gemini 2.5 Pro TTS3.904.144.244.064.264.12
Gemini 3.8 Flash TTS4.174.043.974.084.294.11
Gemini 3.8 Flash-Lite TTS3.884.144.213.964.314.10
ElevenLabs v34.064.313.513.474.153.90
Gemini 2.5 Flash TTS3.244.184.103.613.883.80
Gemini 3.1 Flash TTS3.083.993.913.433.603.60
Higgs Audio v2 (3B)3.404.033.383.103.473.48
VibeVoice 1.5B †3.353.383.393.193.623.39
† VibeVoice 1.5B: hallucinated speech not in the script appears in some clips.
See model comparison performance for gender pairs
ModelMale-femaleFemale-femaleMale-maleSeparation overall
Gemini 3.8 Flash TTS4.924.173.424.17
ElevenLabs v34.833.024.334.06
Gemini 2.5 Pro TTS4.774.172.773.90
Gemini 3.8 Flash-Lite TTS4.883.463.313.88
Higgs Audio v2 (3B)4.902.293.023.40
VibeVoice 1.5B4.383.062.623.35
Gemini 2.5 Flash TTS4.883.081.753.24
Gemini 3.1 Flash TTS3.792.313.123.08

ElevenLabs v3 scores 4.06 on speaker separation and 3.47 on conversation dynamics, the largest gap between the two of any system we evaluated. Keeping two voices apart and making an exchange sound like two people talking are different capabilities, and a system can be strong at one while trailing on the other.

Why new capabilities need new evaluations

The work began with defining what good multi-speaker dialogue looks like and developing rubrics to measure it. Applying those rubrics across model generations made progress and remaining weaknesses visible, including speaker-separation problems that overall scores could conceal. The resulting evaluation gives researchers a repeatable way to identify priorities for the next iteration and assess whether subsequent changes improve performance.

We help research teams develop bespoke rubrics and provide the infrastructure to run private evaluations. If you are developing a new capability, get in touch to discuss how to define quality and measure progress.

Stay in the loop

Get the latest on empathic AI research, product updates, and company news.

Join the community

Connect with other developers, share projects, and get help from the team.

Join our Discord