EU
Biometric Data under GDPR
The short answer. Voice is not automatically special-category data under the GDPR. Article 4(14) defines biometric data as personal data resulting from specific technical processing of physical, physiological or behavioural characteristics that allow or confirm unique identification, and Article 9(1) prohibits processing of biometric data for the purpose of uniquely identifying a natural person. The test is the purpose, not the file format. The same recording is ordinary personal data in a transcription corpus and special-category data in a speaker verification corpus.
The law
GDPR Article 4(14)
'biometric data' means personal data resulting from specific technical processing relating to the physical, physiological or behavioural characteristics of a natural person, which allow or confirm the unique identification of that natural person, such as facial images or dactyloscopic data.
The definition has two limbs: specific technical processing, and the ability to allow or confirm unique identification. A raw voice recording that nobody processes for identification does not meet it. A voice embedding produced by a speaker recognition model does. This is why the same audio can be in and out of the definition depending on the pipeline it enters, and why the label on the dataset is less informative than the model trained on it.
GDPR Article 9(1)
Processing of personal data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, and the processing of genetic data, biometric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person's sex life or sexual orientation shall be prohibited.
The prohibition is not on biometric data as such. It is on biometric data processed for the purpose of uniquely identifying a natural person. That phrase does the work. A speaker identification or voice authentication model is squarely inside it. Diarization that only marks where one speaker stops and another starts is a different case from diarization that resolves each segment to a persistent identity, and the difference is whether the system can match or name the person.
GDPR Article 9(2)(a)
the data subject has given explicit consent to the processing of those personal data for one or more specified purposes, except where Union or Member State law provide that the prohibition referred to in paragraph 1 may not be lifted by the data subject.
Explicit consent is the exception most commercial buyers rely on, and explicit means more than ordinary consent: a specific statement of the purposes, not a general permission. A recording release that says "research and development" does not name speaker identification, and it does not cover it. If the intended use is identification, the consent wording has to say so.
GDPR Article 9(2)(g)
processing is necessary for reasons of substantial public interest, on the basis of Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.
This exception requires a basis in Union or Member State law. It is not available to a commercial buyer on its own judgment, and it is not something a vendor can elect. If a supplier tells you their biometric processing is covered by substantial public interest, the next question is which statute, and the answer is usually that there is not one.
Who it applies to
Article 9 binds any controller or processor handling the data of people in the EU, wherever the company is established, through the territorial rule in Article 3. Two consequences follow for a dataset purchase.
First, the analysis follows the speakers, not the buyer. A corpus recorded in the EU and delivered to a US company is in scope, and the fact that the model is trained outside Europe does not change it.
Second, the special-category character attaches to the processing purpose, which means the same file can move in and out of Article 9 depending on what you do with it. If the intended use is identity verification, authentication, or voice matching, you are in the special-category regime and need an exception from Article 9(2). If the intended use is transcription, translation or synthesis, the analysis usually stays in ordinary personal data — provided speaker identity labels are stripped and the model is not used to identify anyone.
A practical note that catches buyers out: biometric templates are treated as special-category data whether they are stored as an audio sample or as a mathematical representation. Storing a speaker embedding instead of the recording does not move the processing out of Article 9.
What it costs to get wrong
Infringements of Article 9 sit in the higher tier of GDPR Article 83(5): administrative fines up to 20 million euros, or in the case of an undertaking, up to 4% of the total worldwide annual turnover of the preceding financial year, whichever is higher. That tier also covers Articles 5, 6 and 7 and the Chapter V transfer rules, which are the provisions most often infringed at the same time.
Alongside the fine there is Article 82, which gives data subjects a right to compensation for material and non-material damage. National courts have diverged on whether a loss of control over personal data is by itself a compensable harm, and the case law is still developing.
The immediate commercial cost arrives earlier than either. A corpus that was sold for transcription but is used for identification has no valid exception for that use, which means the dataset cannot lawfully support the product it was bought for, whatever was paid for it. The fix is either new consent or a different corpus, and neither is quick.
How to comply when you are buying data
The purpose test is answered on the buyer's side, so the first step is writing down what the data is for before the purchase is scoped.
- State the intended use precisely. "Speaker identification and authentication" triggers Article 9. "Transcription" usually does not. One corpus can serve the second and not the first, and the price should reflect which one you are buying.
- Ask whether the dataset includes speaker embeddings, voiceprints, or identity labels. Those artifacts are what turn an ordinary corpus into a special-category one, and they are usually removable.
- If the use is identification, require explicit consent wording that names it, and read the form rather than the summary. Purposes stated as "data research" or "product improvement" do not cover it.
- Check that the consent also covers the downstream steps: model training, distribution of the model, and sublicensing to third parties. Each is a separate purpose under Article 5(1)(b).
- If Article 9(2)(g) is claimed, ask for the statutory basis in writing. Treat an answer without a citation as a no.
- Keep the DPA and the impact assessment together. Large-scale processing of special-category data is one of the mandatory triggers for a DPIA under Article 35(3)(b).
Related compliance topics
-
Illinois BIPA
bipa
-
Data Protection Impact Assessment
dpia
-
Data Processing Agreement
data processing agreement
-
Vendor Due Diligence Checklist
vendor due diligence checklist
-
Voice Anonymization
voice anonymization
-
Anonymisation vs Pseudonymisation
anonymisation vs pseudonymisation
Not legal advice
We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.
Sourcing data under Biometric Data under GDPR?
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.