What is Speaker Identification?

Identifying who is speaking from the voice itself. Also called voice biometrics.

Training a speaker identification model needs many recordings of the same person, not one recording of many people. The speaker distribution in the dataset sets the ceiling on what the model can do.

The most common collection mistake is too small a speaker pool — a few dozen people recorded for hundreds of hours teaches the model those few dozen voices and nothing else.

Related terms

  • Diarization

    Marking which speaker said which segment in a multi-speaker recording.

  • Annotation

    Attaching machine-readable labels to raw data — a transcript, an intent, a speaker identity.

Keep reading

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com