What is Sample Rate?
How many times per second audio is captured, in hertz (Hz). Speech data is usually delivered at 16kHz or 48kHz.
16 kHz is the common standard for speech recognition because it covers the frequency range that matters for the human voice. 48 kHz is mostly for music and professional recording.
Downsampling is possible; upsampling cannot invent information that was never captured. So when in doubt, record at the higher rate.
Related terms
-
Transcription
Writing down what is said in a recording. It is the core annotation task in any speech dataset.
-
SNR
Signal-to-Noise Ratio — how much louder the speech is than the background noise, measured in decibels (dB).
Keep reading
-
Full glossary
Every term we explain, in one list.
-
How to buy training data
Where these terms actually show up, and which ones change a quote.
-
Help center
How a project runs, from specification to delivery.
Submit a sourcing request
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.