Now liveExplore Expression APIs
hume.ai logo

Hume VoiceEQ Platform

Measure how your voice AI sounds to people. Ship the fix. Measure again.

VoiceEQ gives teams one platform to define what good voice performance means, test under realistic conditions, combine automated measurement with human judgment, and understand where models and agents need to improve.

With Hume’s VoiceEQ

  • Reliability
  • Long-form stability
  • Acoustic quality
  • Expression understanding
  • Emotion alignment
  • Expressivity robustness
  • Voice naturalness
  • Problem redirection
  • Interruption handling
  • Speech expression measurement
  • Speaker verification
  • Synthetic-speech detection
Traditional metrics
  • WER
  • Latency
  • Task success
  • Accents
  • Emotional speech
  • Background audio
  • Conversational speech
  • Warmth
  • Authenticity
  • Pacing
  • Prosody match
  • Disfluency handling
  • Barge-in recovery
  • Empathy fit
  • Escalation handling

An end-to-end platform for improving voice AI, not just scoring it.

Hume combines real-world interactions, expression-rich data, expression measurement, and human ground truth to find failures and guide improvement.

01 / Generate

Automatically turn real-world use cases into targeted evaluation suites.

VoiceEQ analyzes your conversation data to identify use cases, generate test cases, and define the scenarios and success criteria that matter.

Explore the products

Your conversation data
ScenariosPersonasConditionsSuccess criteria
Your evaluation suite

Voice-native infrastructure underneath every evaluation.

VoiceEQ brings together expression measurement, audio data pipelines, and human testing and ratings to evaluate what your AI says, how it sounds, and how people experience it.

Expression Measurement

Audio Expression API
600+ dimensions from the audio itself, in 50+ languages, speaker-aware and timestamped. Streams over SIP, WebRTC, or WebSocket.
Video Expression API
Measure 70+ expressions across 60+ faces per frame. Adapt your system in realtime.

Data engine

Prism Pipeline
A configurable audio pipeline in your cloud: clean, segment, transcribe, enrich, organize. Deployable in your VPC on GKE, EKS, or AKS.
Conversation Intelligence
Your calls, clustered into use cases and failure modes.
Speech Data
Speech, facial-expression, text, and interaction datasets.

Human in the loop

User Testing API
Real people test your endpoint, in multiple languages, across targeted personas, scripts, and edge cases.
Human Feedback API
Ground truth ratings from real humans in hours, not days. Multilingual and cross-cultural.

How it connects

Transports

  • WebRTC
  • WebSocket
  • SIP / live phone line
  • Batch audio

Modalities

  • TTS
  • Speech-to-speech
  • ASR
  • Speech understanding

The same infrastructure behind the largest benchmark of its kind.

The Real-World VoiceEQ benchmark evaluates voice AI across the dimensions people actually experience, using private held-out evaluations and human judgment.

Human ratings
1M+
Speech samples
100M+
Voice modalities
4
  • Speech generation (TTS)
  • Speech-to-speech
  • Speech recognition (ASR)
  • Speech understanding
Real-World VoiceEQ
  1. 01Model 0184
  2. 02Model 0279
  3. 03Model 0371
  4. 04Model 0466
  5. 05Model 0558
  6. 06Model 0651
Illustrative, not real data

One platform. Two ways teams use it.

Develop voice models

Build richer data, evaluate checkpoints, diagnose failure patterns, and improve the next model.

Explore for model builders

Deploy voice agents

Choose the right models, configure behavior, test realistic interactions, and improve the next release.

Explore for agent teams

Evaluation built for sound, not transcripts.

Most evaluation systems begin with transcripts. Hume begins with the audio itself, combining proprietary expression measurement, realistic voice interactions, and human judgment to measure dimensions other systems cannot see.

Voice-native

Measures information carried in the audio, not just the transcript.

Human-grounded

Uses trained human judgment where automated metrics fall short.

Built for improvement

Connects evaluation, failure analysis, data, and repeatable regression testing.

Know how your voice AI performs before your users tell you.