Product
Introducing Expression APIs: Give your AI more of what people mean

Human communication carries meaning beyond words. A change in tone, a hesitant delivery, or a surprised expression can change how we understand what someone says, and how we respond.
Consider “I can’t believe it.” After good news, those words might convey excitement. After a setback, they might express frustration. During a playful exchange, they might communicate amusement. When AI applications hear these words in any of these contexts, the words they transcribe stay exactly the same.
For applications built to interact with people, that missing context matters. A voice agent may understand a customer’s request while missing frustration in their delivery. A researcher may know what participants said while losing the expressive detail that helps explain their reactions.
Today, we’re introducing Hume Expression APIs to make those signals available to developers and researchers. Our models turn vocal and facial expression into structured scores that estimate how people are likely to perceive what an expression conveys, with voice expression also measuring speaking styles and vocal qualities. Expression API supports live interactions as well as analysis and tagging of recordings after the fact.
Our models were trained on over 260,000 speech recordings and 370,000 human-rated images of facial expressions. Evaluated against human judgment (the gold standard for expression measurement), our models match or exceed state-of-the-art models on tested measures of vocal and facial expression. Read the full validation post for the methods, results, and model comparisons.
Audio Expression API: Measure vocal expression, delivery, and vocal qualities

Expression API turns a person’s emotional expression, speaking styles, and vocal qualities into structured scores. Developers can use those signals to inform an agent’s next response, identify difficult moments for training and review, or find recordings with the characteristics their model needs to learn from.
Our vocal-expression model measures expression across 414 tags, comprising 23 higher-level emotion categories and 391 finer-grained descriptions directly from audio. These span amusement, anger, interest, joy, sadness, and surprise, alongside more specific expressions such as admiration, frustration, nervous laughter, and relief.
Our voice descriptor model adds 190 characteristics, including raspy, breathy, fast, monotone, whispering, and theatrical. Together, the models describe combinations such as quietly expressed frustration or amusement in a deep, raspy voice, giving developers richer annotations for analysis and training.
Available in 50+ languages, including Arabic, German, English, Spanish, French, Indonesian, Italian, and Japanese.
Video Expression API: Bring facial expression into the picture

Video Expression makes visible signals that people communicate via facial expression available as structured measurements. Developers can use them to inform live interactions, while researchers can tag recordings, locate reactions worth reviewing, and compare expressive responses alongside what participants say.
Our facial expression model measures 48 emotional-expression categories, including admiration, amusement, anger, anxiety, concentration, confusion, excitement, joy, and sympathy. It also measures 27 visible facial descriptions, such as Smile, Frown, and Jaw drop.
Our facial-expression model draws on a human-rated collection with participants from the United States, South Africa, India, China, Venezuela, and Ethiopia.
Expression APIs in action: put expression measurement to work
Use voice and facial-expression scores during live interactions, or analyze and tag recordings afterward. Timestamped results connect scores to the source material, helping teams review relevant moments and select data for evaluation and training. Batch processing is coming soon.
Help AI agents resolve difficult calls and improve customer satisfaction

An AI agent may understand a customer’s request but miss frustration in their delivery. Real-time vocal-expression scores provide context for acknowledging the difficulty, clarifying an answer, or offering a handoff—with the goal of reducing unnecessary escalations and improving customer satisfaction.
After calls, expression tags can help contact-center teams find moments expressing frustration or anger, prioritize QA review, and select examples for coaching human agents or improving AI behavior.
Make market research faster and capture reactions more fully

Transcripts can flatten participants’ reactions, while reviewing hours of recordings takes time. Voice and Visual Expression APIs help researchers locate and compare expressive responses to a product concept, advertisement, or message alongside participants’ stated feedback.
This focuses manual review and preserves context beyond the transcript. Researchers can investigate where words and delivery differ, interpreting those moments through the recording, follow-up questions, and broader study.
Help robots adapt when an interaction breaks down

A robot may continue its instructions when a person appears confused or expresses frustration. Vocal and facial-expression measurements add context for deciding when to pause, repeat, simplify, or request human assistance.
In health and aged-care settings, developers can use these measurements to build and test interactions designed to reduce misunderstandings and help people complete tasks more comfortably, alongside their words and circumstances.
Find targeted data and test whether your model improves
Finding training examples to address an evaluation weakness can require substantial manual review. Expression and voice tags help researchers retrieve recordings with characteristics such as quietly expressed frustration, whispered speech, or a raspy voice. Facial-expression scores similarly organize visual data.
Teams can preprocess data for evaluation by tagging recordings and building targeted test sets. The same annotations help identify underrepresented expressions and select fine-tuning and post-training examples, reducing search time and supporting progress measurement.
Start building with Hume’s Expression API
Talk to Hume about API access and how voice and visual expression can fit your application, research, or model-development workflow.
Keep reading

Building Voice Models Is No Longer a Modeling Problem
What’s changed isn’t just where voice is used, but what it represents. Voice is no longer a feature layered on top of an intelligent system. It’s becoming a foundational modality through which models reason, interact, and are judged by users.
Jeremy Baum / Jan 21, 2026

Octave 2: next-generation multilingual voice AI
Today we’re launching Octave 2, the second generation of our frontier voice AI model for text-to-speech. We just made a preview of Octave 2 available on our platform and through our API.
Oct 1, 2025