Now liveExplore Expression APIs
hume.ai logo
All posts

Article

Behind Hume’s Expression Measurement Models for Face and Voice

hume.ai logo
Hume AI Team
Share
Behind Hume’s Expression Measurement Models for Face and Voice

Editor's note: This post covers the science, development and evaluation of Hume’s state of the art Expression API, that enables developers to enrich their applications or research pipelines with real-time measurement of voice and facial expressions.

Facial expressions and vocal tone are central to how we communicate. They convey emotions such as amusement, interest, frustration, and surprise, often without naming those feelings in words. Understanding these signals is part of how we connect with one another and respond to what others express.

Expression measurement turns these signals into structured scores that can be used to annotate data, evaluate AI outputs, and explore expressive audio and images. It gives us a way to describe the emotional meaning people perceive, compare expressions across recordings, and follow how an interaction changes over time.

Consider someone saying, “I can’t believe it.” After good news, the phrase might sound excited. After a setback, it might convey frustration. During a playful exchange, it might express amusement. The words stay the same, while the voice and facial expression contribute different meanings.

At Hume, our work on expression measurement is grounded in emotion science. Our research has explored the range of meanings people associate with facial and vocal expressions, including nuanced and blended expressions. This scientific perspective guides our approach: measure expression through a detailed vocabulary and evaluate those measurements against human judgments. Read more about the science behind Hume’s expression measurements.

An expression can communicate several emotional qualities at once, such as amusement blended with disbelief. Our speech emotion model captures this detail through separate scores for 414 tags: 23 core emotions and 391 finer-grained descriptions. These include amusement, anger, interest, joy, sadness, and surprise, alongside more specific expressions such as admiration, frustration, nervous laughter, and relief.

The voice descriptor model measures 190 characteristics across 27 families, describing qualities such as raspy, breathy, fast, monotone, whispering, and theatrical. These measurements complement emotional expression, allowing us to describe combinations such as frustration expressed quietly or amusement in a deep, raspy voice. Our speech models support Arabic, Bengali, German, English, Spanish, French, Hebrew, Indonesian, Italian, Japanese, Korean, Polish, Portuguese, Russian, Turkish, and Vietnamese.

Our facial expression model recognizes emotional expression across 48 Hume categories, including admiration, amusement, anger, anxiety, concentration, confusion, excitement, joy, and sympathy. It captures nuanced expressions and distinguishes positive from negative surprise. The model also measures 27 visible facial descriptions, such as Smile, Frown, and Jaw drop, providing a complementary view of an expression’s visible features.

Data and human ratings behind expression measurement

For facial expression, our training data includes more than 370,000 images from our human-rated FaceMimic collection. Across the collection, more than 13,000 raters contributed more than one million emotion ratings, describing both the emotional meaning of an expression and its intensity. Participants photograph themselves mimicking a reference expression, then rate their own and others’ expressions. The collection includes participants from the United States, South Africa, India, China, Venezuela, and Ethiopia.

Our speech emotion and voice descriptor models share a speech training resource with more than 260,000 recordings. Its emotion-annotated portion includes more than 210,000 recordings, representing more than 270 hours of mostly performed speech, with multiple emotion annotations for most clips. The resource captures both emotional expression and vocal characteristics, supporting measurements of what a voice conveys and how it sounds.

Evaluating agreement with human perception

We evaluated our voice models using our Hume Feedback API, collecting structured judgments from listeners, including native speakers for multilingual studies. We asked whether the models’ descriptions agreed with human ratings and whether higher scores identified stronger examples of a particular expression or vocal characteristic.

For each model, we begin with evaluations on our internal human-rated data, then examine results on public datasets. These studies assess how well the models recognize expression and capture the strength of the qualities people perceive.

Within each study, the compared systems were assessed on the same evaluation samples.

Speech Emotion Recognition Results

We compared 11 systems on the same set of more than 9,000 recordings from 32 public datasets, including Gemini 3.8 Flash and specialized systems such as EmotionThinker, EmpathicInsight, and VoiceCLAP. All systems made zero-shot predictions: we applied the models directly, without task-specific fine-tuning for this benchmark. Our speech emotion model achieved the highest overall AUC, with statistically significant advantages over each of the other systems shown. Its advantage over Gemini 3.8 Flash also remained statistically significant when we evaluated the full test sets of more than 43,000 recordings and checked different combinations of datasets.

We use one metric throughout the public benchmark tables: emotion-ranking score, also called AUC. Imagine taking two recordings: one labeled sad and one without sadness. The score measures how reliably the model gives the sad recording a higher sadness score, averaged across the evaluated emotions. Higher is better: 50 represents chance-level ranking and 100 represents perfect ranking. This measures how well recordings are ordered, rather than the percentage of labels guessed correctly. The overall table shows all 11 systems ordered by this score. EmpathicInsight and VoiceCLAP share the same rounded score.

Emotion-ranking score (AUC)

System AUC
1Hume Expression API 77.1
2Gemini 3.8 Flash 73.0
3EmotionThinker 71.0
4EmpathicInsight 67.1
4VoiceCLAP 67.1
6gpt-realtime 64.2
7MOSS-Audio 60.5
8Qwen2.5-Omni 59.9
9Grok 57.1
10MERaLiON 54.8
11Phi-4 51.1

Bars run from 50 (chance) to 100 (perfect). Best value in accent.

The table below groups all 32 public datasets by spoken language and uses the same emotion-ranking score as the overall comparison. Each dataset receives equal weight within its language group, and every model is evaluated on the same recordings. The final row gives the overall mean across all 32 datasets, with each dataset weighted equally. Bold values mark the highest displayed score in each row, including ties at the shown precision.

Scores are averaged across the public datasets for each language, with each dataset given equal weight.

Emotion-ranking score by language

Language Dataset
count
Hume Gemini
3.8 Flash
Emotion
Thinker
Empathic
Insight
Voice
CLAP
English1077.270.473.865.467.5
Bengali476.169.967.468.163.4
Polish269.956.358.861.962.3
Italian272.864.960.962.162.3
Indonesian280.380.471.368.763.1
French279.483.477.077.474.7
Russian273.381.467.462.762.2
Urdu175.576.068.762.673.7
Turkish173.358.168.963.455.8
Persian188.990.883.179.080.6
German178.472.977.469.276.2
Thai177.585.675.167.370.0
English and French174.980.165.372.171.6
Portuguese185.383.981.668.474.8
Amharic186.273.673.173.466.4
Mean3277.173.071.067.167.1

Emotion-ranking score (AUC) by spoken language. Bold marks the highest displayed score in each row, including ties at the shown precision.

We then compared our speech emotion model with Gemini on internal human ratings collected through our Hume Feedback API. This evaluation covers finer-grained emotional expressions across 16 languages, excluding the broad parent emotion categories. We use the same emotion-ranking metric as in the public benchmark: AUC, shown on a 0-100 scale, with higher scores indicating better ranking. The mean gives each language equal weight.

On our internal human-rated evaluation of finer-grained emotional expressions across 16 languages, our model performed comparably to Gemini.

System Emotion-ranking score (AUC)
Hume Expression API58.9
Gemini 3.8 Flash59.6

Internal human-rated evaluation of finer-grained emotional expressions across 16 languages, with each language weighted equally.

On an English evaluation of emotion labels named by multiple annotators, our model achieved a statistically significant emotion-ranking advantage over Gemini 3.8 Flash.

Voice Descriptor Characteristics Results

Voice characteristics give us another way to describe and organize speech. Someone searching an audio collection might want to find raspy voices, fast delivery, or theatrical expression. Our Hume Expression API assigns scores to these qualities, helping bring relevant recordings to the top.

To evaluate how well this works, we use a descriptor-ranking score, measured by average precision. For example, when recordings are sorted by their raspiness scores, those that listeners described as raspy should appear near the top. We average this measure across the evaluated characteristics and report it on a 0–100 scale, where higher scores indicate better ranking.

Using human ratings collected through our Hume Feedback API, we compared our voice descriptor model with Gemini 3.8 Flash, GPT-Audio-1.5, Grok Voice Think Fast 2.0, and specialist models on an evaluation set of more than 2,000 recordings across 16 languages. Listeners assessed whether each vocal characteristic was present and how strongly it came across. Alongside the descriptor-ranking score, we report intensity-ranking agreement: how closely a model follows listeners’ judgments of which recordings express a characteristic more strongly. This is measured by Spearman correlation, with higher scores indicating closer agreement.

Voice descriptor characteristics results
System Descriptor-ranking
score (mAP, 0–100)
Intensity-ranking
agreement (Spearman)
1Hume Expression API84.00.380
2Gemini 3.8 Flash83.40.365
3ParaSpeechCLAP Intrinsic75.20.030
3GPT-Audio-1.575.20.270
5LAION-CLAP74.90.144
6Vox-Profile WavLM73.70.033
7ParaSpeechCLAP Combined72.0−0.020
8Vox-Profile Whisper70.1−0.020
9Grok Voice Think Fast 2.068.70.150

Nine systems, ranked by descriptor-ranking score. Bold marks the highest score in each column.

Facial Expression Recognition Results

Facial expressions can convey several emotional qualities at once. A smile might communicate amusement, relief, or excitement, depending on the expression. Our facial expression model assigns scores across 48 emotional categories, alongside descriptions of visible features such as a smile, frown, or dropped jaw.

We use the same expression-ranking score, AUC, as in the voice sections. For amusement, for example, a strong model gives higher amusement scores to faces viewers labeled amused than to faces they did not. We average this comparison across the evaluated emotions and report it on the same 0–100 scale.

We first compared our model with 25 other models on EmoNet-Face HQ, a public benchmark of synthetic faces rated by psychology experts. All models were evaluated on the same set of about 1,200 images across 40 emotions. The comparison included Gemini, GPT, Grok, Qwen, and specialist facial emotion models.

Hume achieved the highest expression-ranking score, with statistically significant advantages over all 25 comparators, including Gemini 3.8 Flash. The table shows Hume and the four strongest comparator families, with the highest-scoring model selected from each family.

System Expression-ranking score (AUC)
1Hume Expression API76.9
2Gemini 3.5 Flash71.3
3Qwen3-VL-32B68.4
4GPT-5.567.0
5Gemma 4 26B-A4B64.0

Highest-scoring model from each comparator family. Bold marks the highest score.

Gemini 3.8 Flash scored 68.0 in the same evaluation. These results measure how reliably the models distinguish faces with a human-rated emotional quality from faces without it.

We then examined how the models agreed with human judgments on our internal FaceMimic evaluation. This includes about 2,100 photographs of individual expressions, each rated by at least two people, across the 48 Hume emotions. Hume, Gemini 3.8 Flash, and Qwen3.8-27B were evaluated on the same images using the same expression-ranking metric.

System Expression-ranking score (AUC)
Hume Expression API77.9
Gemini 3.8 Flash69.5
Qwen3.8-27B61.1

Internal FaceMimic evaluation: ~2,100 photographs of individual expressions, each rated by at least two people across the 48 Hume emotions. Bold marks the highest score.

Our model achieved statistically significantly higher expression-ranking scores than both Gemini and Qwen in this internal evaluation. Together, the public benchmark and internal human ratings show strong recognition of the emotional qualities people perceive in facial expressions.

We also compared Hume with dedicated facial emotion recognition models on the same held-out FaceMimic images. To make this comparison fair at the evaluation level, each pair received identical images and was scored against the same averaged human ratings, using only emotions both models could represent. The emotion mappings and image crop rules were fixed before scoring, and emotions without an aligned counterpart were excluded from both models.

Expression-ranking score (AUC; higher is better) on shared emotions

Specialist model Shared
emotions
Hume Specialist
model
EmotiEffLib default ENet-B0681.170.9
Empathic Insight Face Large2478.169.0

Expression-ranking score (AUC; higher is better) on shared emotions. Each row uses its own shared emotion vocabulary, so scores should be compared within a row rather than against the 40- or 48-emotion tables above. Bold marks the higher score in each row.

Put expression measurement to work

These measurements can become part of an existing workflow: adding expression labels to a dataset, comparing the delivery of generated speech, or helping reviewers find relevant moments in recorded interactions. Our human-rated evaluations provide evidence for choosing the measurements that fit each task.

Contact Hume for Expression Measurement API access and to discuss how it can support your data annotation, model evaluation, or product workflow.


Keep reading

Evaluating Google’s multi-speaker TTS: A case study in why private evaluations matter

Text-to-speech systems were originally built to read text aloud in a single voice. As voice AI expands into audiobooks, game dialogue, and advertising, these systems are being asked to do more: generate conversations between multiple speakers. That takes more than generating two distinct voices and stitching their lines together. The speakers need to sound like they are responding to one another, with natural timing, changes in tone, and smooth handoffs. Each voice must remain distinct and consistent while contributing to a believable conversation.

Sharath Rao, Kimberly Lo, Alice Baird / Sep 29, 2026

Newly-released Google's Gemini 3.8 Flash TTS tops Hume's Real-World VoiceEQ leaderboard

Voice models are advancing quickly, bringing new capabilities to a growing range of use cases. Evaluating them requires assessing a broader range of qualities—and whether models can sustain their performance over longer conversations and passages. Hume’s Real World VoiceEQ benchmark takes this broader view of performance, measuring the human qualities of speech, including tone, expression, naturalness, speaker identity, and reliability, across real-world conditions and use cases. We recently extended our evaluations to cover voice controllability and voice replication.

Alice Baird / Sep 24, 2026