Now liveExplore Expression APIs
hume.ai logo

Our Science

The science behind emotionally intelligent voice AI

Hume's work is based on over a decade of research into human emotion and expression together with our experience building voice models. We turn that foundation into the data, measurements, and feedback researchers need to develop voice AI that listens, speaks, and responds more appropriately.

AmusementAweDistressContemplation
Expression embedding, judged by people

From emotion theory to computational science

David Hume’s theories of emotion and Charles Darwin’s observations of expression laid early groundwork. Experimental and cross-cultural studies in the 1960s strengthened the field’s scientific foundations.

Today, Hume AI combines computational methods with large-scale human studies to turn that science into data, measurements, and feedback for voice AI development.

Learn more

A history of emotion science

  1. 1739

    David Hume

    Theorizes that emotions drive choice and well-being.

  2. 1872

    Charles Darwin

    Observations of facial, bodily, and vocal expression.

  3. 1960s

    Emotion research

    Researchers undertake experimental and cross-cultural studies on emotional expression.

  4. Today

    Hume AI

    Hume’s scientists use data-driven methods, computational techniques, and large-scale human studies to understand how people express and perceive emotion across voice, face, and language.

Voice carries more than words

Meaning comes from the sounds we make, how we deliver words, and how listeners interpret them. Research by Hume and our collaborators helps explain these layers and provides a scientific foundation for building voice AI.

  • Vocal bursts

    Brief, nonverbal sounds such as laughs, sighs, cries, and gasps convey a range of emotional meanings. U.S. English-speaking listeners distinguished at least 24 meanings in these sounds.

    Read the study(opens in a new tab)
  • Prosody

    Prosody is the pitch, pace, rhythm, loudness, and emphasis in speech. With words held constant, English-speaking listeners in the U.S. and India distinguished at least 12 shared emotional meanings.

    Read the study(opens in a new tab)
  • Cross-cultural perception

    People across cultures interpret vocal expression in both shared and distinct ways. A study across five countries found that vocal bursts conveyed both shared and culture-specific meanings.

    Read the study(opens in a new tab)

Mapping the meanings of expression

Semantic Space Theory combines large-scale human studies with statistical modeling to map the emotional meanings people perceive in expression, and how those meanings blend.

Hume uses this framework to define expression categories, annotate data, and evaluate voice models.

Learn more about Semantic Space theory
Explore the maps
Speech Prosody · Model outputsOriginal map

Loading Hume’s semantic-space map…

Explore the research(opens in a new tab)

We made the science usable for model development

Building expressive voice models required us to translate studies of human perception into engineering decisions: what to represent, which data to collect, how to annotate it, and how to tell whether a model improved.

Represent expression

Our detailed taxonomy represents combinations such as quiet anger or restrained excitement that simpler labels can miss.

Explore our 600+ tags
Expression taxonomy

600+Tags

Emotion categories
Afraid, Angry, Amused, Joyful, Disgusted, Excited, Distressed, Sad, Surprised
Speaking styles
Whisper, Calm, Excited, Nervous, Assertive
Vocal qualities
Pace, Warmth, Energy, Clarity, Hesitation

Build the right data

Our cross-product sampling research explores expressive combinations that ordinary filtering can miss, giving models richer examples to learn from.

Explore cross-product sampling
Cross-product sampling
  • Quiet + angerExpression × delivery
  • Quiet + excitementExpression × delivery
  • Forceful + angerExpression × delivery
  • Forceful + excitementExpression × delivery

Conceptual illustration of sampling combinations

Design useful measurements

Our Real-World VoiceEQ benchmark examines voice performance through human-grounded evaluations and compares automated judgments with human ratings.

Read the benchmark whitepaper(opens in a new tab)
Research preprint
  • Speech generation
  • Conversation
  • Understanding
  • Transcription
  1. 01Defined tasks
  2. 02Listening rubrics
  3. 03Scoring methods

Connect feedback to model development

For a major frontier lab, we developed bespoke evaluations of multi-speaker dialogue to identify priorities for the next model iteration.

Read the frontier-lab case study
Compare model generationsIllustrative comparison

Bring your next voice research question to Hume

Work with researchers and engineers who connect the science of human expression to your model-development goals. We help define the right evaluations, identify improvement priorities, and measure progress across model generations.