EU

Is Voice Biometric Data?

The short answer. It depends on what you do with it, not on what it is. Under GDPR, voice is biometric data only when it is processed through specific technical means for the purpose of uniquely identifying a person. Training an ASR model on voice is generally not that, because the purpose is transcription. Building a speaker identification or verification system is, because unique identification is the whole point. That distinction decides whether Article 9 applies, and it is usually the buyer's intended use rather than the supplier's collection method that settles it.

The law

GDPR Article 4(14)

"biometric data" means personal data resulting from specific technical processing relating to the physical, physiological or behavioural characteristics of a natural person, which allow or confirm the unique identification of that natural person, such as facial images or dactyloscopic data.

Three conditions have to hold at once: the data results from specific technical processing, it relates to physical or behavioral characteristics, and it allows or confirms unique identification. Audio used for transcription does not satisfy the third, and unprocessed audio may not satisfy the first.

GDPR Article 9(1)

Processing of personal data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, and the processing of genetic data, biometric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person's sex life or sexual orientation shall be prohibited.

Note the qualifier: the prohibition reaches biometric data for the purpose of uniquely identifying a natural person, not biometric data in general. That is why the same recording can be ordinary personal data in one product and special category data in another.

GDPR Recital 51

The processing of photographs should not systematically be considered to be processing of special categories of personal data as they are covered by the definition of biometric data only when processed through a specific technical means allowing the unique identification or authentication of a natural person.

This is the reasoning that also applies to voice. A photograph is not special category data because it shows a face; it becomes biometric data when the processing is aimed at identifying the person. Substituting voice for photograph gives you the test.

GDPR Article 9(2)(a)

Paragraph 1 shall not apply if one of the following applies: (a) the data subject has given explicit consent to the processing of those personal data for one or more specified purposes, except where Union or Member State law provide that the prohibition referred to in paragraph 1 may not be lifted by the data subject.

Where Article 9 does apply, explicit consent is the usual route. Explicit is a higher standard than ordinary consent: it should be a specific and unambiguous statement of agreement to the biometric processing, given separately. A general consent to "use of my voice" is not a safe candidate.

Who it applies to

The analysis follows the purpose of the processing, which means it can change hands. A supplier can deliver a corpus that is not biometric data in their own hands, and the buyer can turn it into biometric processing by using it to train an identification model. The buyer does not escape Article 9 by pointing out that the supplier did not ask for explicit consent.

Two uses are easy to confuse and land on opposite sides. Identification answers the question "who is this speaker", and it is the paradigm Article 9 case. Verification or authentication answers "is this the person who claims to be here", and Article 4(14) covers both, because the definition says allow or confirm unique identification. A wake-word system that merely detects that speech is present, without attributing it to a person, is neither.

There is a middle case worth flagging. Emotion recognition and speaker diarization do not identify anyone, so they are not automatically Article 9 processing, but emotion recognition from voice engages the EU AI Act transparency duties, and diarization produces a speaker-specific pattern that can become identifying when combined with other data. Neither should be assumed to be ordinary processing without a written reason.

The practical implication for a data deal is that the intended purpose has to be stated. A supplier selling one corpus to an ASR team and a speaker-verification team is selling two different regulatory products, and the consent documentation is not the same for both.

  • ASR, TTS and diarization training: personal data, not special category, provided no identification is attempted.
  • Speaker identification or verification: Article 9 processing, requiring an Article 9(2) condition.
  • Emotion recognition: not Article 9 by default, but subject to its own disclosure duties under the EU AI Act.

What it costs to get wrong

Where Article 9 applies and no Article 9(2) condition exists, the processing is prohibited, and the fine tier is the higher one: up to 20 million euros or 4% of total worldwide annual turnover under Article 83(5), because Article 83(5)(a) covers the conditions for processing under Articles 5, 6, 7 and 9.

There is also a compensation route under Article 82. Biometric processing claims are attractive to bring because the harm is characterised as inherent rather than needing to be proved item by item, and collective actions are available in several member states.

The consequence that shapes most deals is contractual. EU enterprise customers ask whether the dataset is biometric data and treat an unclear answer as a stop. A supplier who says "voice is always biometric" and a supplier who says "voice is never biometric" are both giving answers that counsel cannot work with, and both cost the deal.

For a buyer who has already built the model, the position is worse than a fine. A speaker identification system trained on data collected without an Article 9(2) condition cannot be made lawful by adding a document, so the remediation is rebuilding the model on a correctly documented corpus.

How to comply when you are buying data

The decisive question is what you intend the data to train. Answer that in writing before you buy, because the answer determines which consent documentation you need and whether the corpus you are looking at is the right one.

One question decides most of this page, and it is a question about your product rather than about the dataset: will the system attribute a voice to a person? If the answer is no, the corpus is ordinary personal data. If the answer is yes, it is Article 9 processing and the consent record has to say so. Buyers who cannot answer it should plan on the stricter position.

Where a project touches identification, we document the intended purpose and the consent basis explicitly rather than describing the dataset as generic voice data, and we decline work where the speakers were not told that their voice could be used to identify them. We are a sourcing company rather than a law firm, so the legal assessment of whether a particular use is Article 9 processing remains one for your counsel.

  • State the intended purpose in the contract. "Voice data" is not a purpose. Whether the corpus will train recognition, synthesis, diarization or identification is the fact that decides the analysis.
  • Ask whether speaker identifiers are shipped with the files. A corpus that carries per-speaker labels is the raw material of identification even if that is not why you bought it.
  • If you intend biometric use, ask for explicit consent documentation as a separate item, and check that it names identification as a purpose.
  • Ask whether the supplier collected the data for a different purpose originally. A corpus repurposed from transcription to identification has a consent gap that cannot be closed after the fact.
  • Do not rely on anonymisation to remove the Article 9 question. Voice is the identifier, so anonymisation strong enough to defeat identification usually destroys the data's training value as well.

How we handle consent and licensing →

Related compliance topics

Not legal advice

We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.

Sourcing data under Is Voice Biometric Data??

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com