What is ASR?

Automatic Speech Recognition — the task of turning recorded speech into text.

An ASR system needs a large volume of paired audio and text to work at all. That pairing is the most basic form a speech dataset takes: one recording, one verbatim transcript.

Three things worth asking before buying ASR training data: how many distinct speakers are in it (not how many hours), what the recording conditions were, and whether the transcript convention follows pronunciation or standard orthography.

Related terms

  • TTS

    Text-to-Speech — the task of generating spoken audio from written text.

  • WER

    Word Error Rate — the standard accuracy metric for speech recognition. Lower is better.

  • Transcription

    Writing down what is said in a recording. It is the core annotation task in any speech dataset.

Keep reading

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com