Speech Data
Richly annotated data for training AI.
Access speech and interaction datasets grounded in Hume’s research. Explore curated collections or work with our researchers to find the data your models need.
Explore the full range of human expression.
Hume offers expression-rich, off-the-shelf and custom speech datasets across languages, cultures, and contexts.
Conversational Data
Single Channel
Multichannel
Dual-channel, 3-Channel
Multilingual
50+ languages
Cross-lingual
Interruption Data
Real world Data
Call Center
Vocal Burst Data
Hard Words
Voice Creation
Human Rated Conversational Data
Data that fits your research workflow.
Explore samples, license the datasets you need, and access training data programmatically through an API that integrates into your existing research pipeline.
Build a dataset around your model’s needs.
Start with an existing collection or work with Hume to scope a custom dataset. Select for the languages, contexts, recording conditions, and target behaviors that matter to your research.
- Starting point
- Multilingual conversation
- Language
- Japanese, native speakers
- Context
- Healthcare intake calls
- Speaking style
- Calm, measured delivery
- Participants
- Two-person dialogue, patient and clinician
- Recording
- Phone-line and headset audio
- Target behaviors
- InterruptionsClarifying questionsTurn-taking
Used for
Give your models more to learn from.
Voice generation
Train speech models for expressive delivery, distinctive voices, and reliable pronunciation.
Speech Understanding
Develop models that meet human expectations with expressive voices and adaptation.
Conversational AI
Train for turn-taking, interruptions, and the reactions that shape an interaction.
Domain adaptation
Fine-tune with conversations relevant to healthcare, education, customer service, and other settings.
Why Our Datasets
World-class data for pre-training and fine-tuning your own models, backed by years of scientific research.
Ethically sourced
- All data collected with informed consent and rigorous privacy protections.
Globally diverse
- Representative samples across cultures, ages, genders, and demographics.
Expert annotated
- Labeled by trained researchers using validated scientific frameworks.
Research ready
- Clean, structured formats optimized for modern ML pipelines.
Find the data for your next breakthrough.
Tell us what you’re building, the capabilities you want to improve, and the examples you need. Explore relevant samples with our researchers and discuss an existing collection or custom data project.