Now liveExplore Expression APIs
hume.ai logo

Human Feedback API

Human ratings for better voice AI.

Evaluate audio, compare samples, and test live conversations with real people. Bring structured human feedback into your development pipeline.

Contact Research

Ask the questions that matter to your model.

Launch your evaluation with a simple API call. Define your questions, instructions, and participant criteria using Hume’s templates or your own study design.

Study setupIllustrative
Audio files
sample_014.wavsample_015.wav
Questions
How natural does this voice sound?
Instructions
Listen to the full clip before rating.
Metadata
model: checkpoint_b
Rater criteria
Native English speakerScreened
Study config
template: naturalness
POST /studies200 OK
{
  "id": "<study_id>",
  "status": "COMPLETED",
  "path": "gs://your-bucket/tts-eval",
  "responses": {
    "gs://your-bucket/tts-eval/clip_001.wav": {
      "naturalness": {
        "values": [2, 4, 4],
        "metrics": { "mean": 3.33, "std": 0.94 }
      },
      "voice_match": {
        "values": [4, 3, 4],
        "metrics": { "mean": 3.67, "std": 0.47 }
      }
    }
  },
  "participants": { "...": "..." }
}

Human feedback in hours, not weeks.

Rate a clip, compare alternatives, or test a conversation. Results are linked to your samples and ready for evaluation.

Single-sample rating

People listen to an individual recording and answer questions you define.

Audio can appear alongside transcripts, tags, prompts, and speaker information. Written responses can provide transcriptions, qualitative comments, and error analysis.

Evaluation areas
  • Transcript accuracy
  • Voice likability
  • Emotional expression
  • Audio quality
  • Naturalness
  • Pronunciation
  • Speaker trustworthiness
  • Role fit

Used for

Understand what people hear and prefer.

  • Model development

    Compare checkpoints to understand whether listeners prefer the changes.

  • Voice selection

    Assess whether a voice fits its intended role, from teaching to narration.

  • Speech quality

    Evaluate naturalness, pronunciation, emotional expression, and listening comfort.

  • Conversational AI

    Learn how people experience a model’s responsiveness, helpfulness, and conversational flow.

Human judgment, backed by voice AI research.

Vetted participants.
Screening, fraud detection, and performance-based incentives support thoughtful, reliable feedback.
Human study experience at scale.
We collected more than 1M human ratings during the development of Real World VoiceEQ.
Global participant network.
Reach participants across countries, languages, and demographics to understand how different audiences experience your voice AI.
Explore the research

Build your next evaluation around human feedback.

Talk with Hume’s research team about your models, evaluation questions, and integration needs.