What is Speaker Identification?
Identifying who is speaking from the voice itself. Also called voice biometrics.
Training a speaker identification model needs many recordings of the same person, not one recording of many people. The speaker distribution in the dataset sets the ceiling on what the model can do.
The most common collection mistake is too small a speaker pool — a few dozen people recorded for hundreds of hours teaches the model those few dozen voices and nothing else.
Related terms
-
Diarization
Marking which speaker said which segment in a multi-speaker recording.
-
Annotation
Attaching machine-readable labels to raw data — a transcript, an intent, a speaker identity.
Keep reading
-
Full glossary
Every term we explain, in one list.
-
How to buy training data
Where these terms actually show up, and which ones change a quote.
-
Help center
How a project runs, from specification to delivery.
Submit a sourcing request
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.