Hume VoiceEQ Platform
Measure how your voice AI sounds to people. Ship the fix. Measure again.
VoiceEQ gives teams one platform to define what good voice performance means, test under realistic conditions, combine automated measurement with human judgment, and understand where models and agents need to improve.
With Hume’s VoiceEQ
- Reliability
- Long-form stability
- Acoustic quality
- Expression understanding
- Emotion alignment
- Expressivity robustness
- Voice naturalness
- Problem redirection
- Interruption handling
- Speech expression measurement
- Speaker verification
- Synthetic-speech detection
- WER
- Latency
- Task success
- Accents
- Emotional speech
- Background audio
- Conversational speech
- Warmth
- Authenticity
- Pacing
- Prosody match
- Disfluency handling
- Barge-in recovery
- Empathy fit
- Escalation handling
An end-to-end platform for improving voice AI, not just scoring it.
Hume combines real-world interactions, expression-rich data, expression measurement, and human ground truth to find failures and guide improvement.
01 / Generate
Automatically turn real-world use cases into targeted evaluation suites.
VoiceEQ analyzes your conversation data to identify use cases, generate test cases, and define the scenarios and success criteria that matter.
Explore the products
Voice-native infrastructure underneath every evaluation.
VoiceEQ brings together expression measurement, audio data pipelines, and human testing and ratings to evaluate what your AI says, how it sounds, and how people experience it.
Expression Measurement
- Audio Expression API
- 600+ dimensions from the audio itself, in 50+ languages, speaker-aware and timestamped. Streams over SIP, WebRTC, or WebSocket.
- Video Expression API
- Measure 70+ expressions across 60+ faces per frame. Adapt your system in realtime.
Data engine
- Prism Pipeline
- A configurable audio pipeline in your cloud: clean, segment, transcribe, enrich, organize. Deployable in your VPC on GKE, EKS, or AKS.
- Conversation Intelligence
- Your calls, clustered into use cases and failure modes.
- Speech Data
- Speech, facial-expression, text, and interaction datasets.
Human in the loop
- User Testing API
- Real people test your endpoint, in multiple languages, across targeted personas, scripts, and edge cases.
- Human Feedback API
- Ground truth ratings from real humans in hours, not days. Multilingual and cross-cultural.
How it connects
Transports
- WebRTC
- WebSocket
- SIP / live phone line
- Batch audio
Modalities
- TTS
- Speech-to-speech
- ASR
- Speech understanding
The same infrastructure behind the largest benchmark of its kind.
The Real-World VoiceEQ benchmark evaluates voice AI across the dimensions people actually experience, using private held-out evaluations and human judgment.
- Human ratings
- 1M+
- Speech samples
- 100M+
- Voice modalities
- 4
- Speech generation (TTS)
- Speech-to-speech
- Speech recognition (ASR)
- Speech understanding
- 01Model 0184
- 02Model 0279
- 03Model 0371
- 04Model 0466
- 05Model 0558
- 06Model 0651
One platform. Two ways teams use it.
Develop voice models
Build richer data, evaluate checkpoints, diagnose failure patterns, and improve the next model.
Explore for model buildersDeploy voice agents
Choose the right models, configure behavior, test realistic interactions, and improve the next release.
Explore for agent teamsEvaluation built for sound, not transcripts.
Most evaluation systems begin with transcripts. Hume begins with the audio itself, combining proprietary expression measurement, realistic voice interactions, and human judgment to measure dimensions other systems cannot see.
Voice-native
- Measures information carried in the audio, not just the transcript.
Human-grounded
- Uses trained human judgment where automated metrics fall short.
Built for improvement
- Connects evaluation, failure analysis, data, and repeatable regression testing.