What is ASR?
Automatic Speech Recognition — the task of turning recorded speech into text.
An ASR system needs a large volume of paired audio and text to work at all. That pairing is the most basic form a speech dataset takes: one recording, one verbatim transcript.
Three things worth asking before buying ASR training data: how many distinct speakers are in it (not how many hours), what the recording conditions were, and whether the transcript convention follows pronunciation or standard orthography.
Related terms
-
TTS
Text-to-Speech — the task of generating spoken audio from written text.
-
WER
Word Error Rate — the standard accuracy metric for speech recognition. Lower is better.
-
Transcription
Writing down what is said in a recording. It is the core annotation task in any speech dataset.
Keep reading
-
Full glossary
Every term we explain, in one list.
-
How to buy training data
Where these terms actually show up, and which ones change a quote.
-
Help center
How a project runs, from specification to delivery.
Submit a sourcing request
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.