What is Biometric Data?

Biometric Data — data derived from a person's physical or behavioral characteristics, such as voice or face, which most privacy laws treat as a special and more tightly restricted category.

Biometric data is information about a person's body or behavior used to identify them. Fingerprints, facial images, iris scans, voice prints and gait are the common examples. What makes the category distinct is that the data is the person: unlike a password, it cannot be changed after a breach.

Because of that, most privacy regimes single it out for stricter treatment. The GDPR treats biometric data used for unique identification as a special category, requiring a higher legal basis than ordinary personal data. Several US states — Illinois, Texas and Washington among them — have dedicated biometric statutes with their own consent requirements, and Illinois' law carries a private right of action that has produced a long line of litigation.

Voice sits in an awkward position in this landscape. A speech recording is not automatically biometric data — someone reading a script is ordinary personal data — but a voiceprint, or any processing whose purpose is to identify a person by voice, generally is. The distinction turns on the purpose of the processing rather than the file format, which is why the same recording can be ordinary data in one project and regulated biometric data in another.

For AI training the practical questions are whether the intended use involves identification, and whether the applicable jurisdiction has a biometric statute. A dataset collected for text-to-speech synthesis and a dataset collected to build speaker recognition are made of similar audio and carry very different legal exposure.

Related terms

  • PII

    Personally Identifiable Information — anything that can identify a specific individual, directly or indirectly.

  • Consent

    The speaker knows how their recording will be used, and has signed to say so.

  • Anonymization

    Anonymization — removing or altering identifying information so a person can no longer be identified, in principle irreversibly, which is what separates it from pseudonymization.

Keep reading

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com