EU

Voice Data under GDPR

The short answer. A voice recording of an identifiable person is personal data, and holding it is processing. That means you need a lawful basis, and for AI training the realistic basis is consent or legitimate interests. You also need the speaker to have been told the specific purposes, because Article 13 puts that duty on whoever collected the data, and a buyer inherits the gap if it was never done. Where the voice is used to identify people, Article 9 raises the bar further. Check the basis, the notice and the transfer mechanism before you take delivery.

The law

GDPR Article 4(1)

"personal data" means any information relating to an identified or identifiable natural person ("data subject"); an identifiable natural person is one who can be identified, directly or indirectly, in particular by reference to an identifier such as a name, an identification number, location data, an online identifier or to one or more factors specific to the physical, physiological, genetic, mental, economic, cultural or social identity of that natural person.

A recording of a speaker who can be identified is personal data, and a voice is a factor specific to their physical and physiological identity.

GDPR Article 4(14)

"biometric data" means personal data resulting from specific technical processing relating to the physical, physiological or behavioural characteristics of a natural person, which allow or confirm the unique identification of that natural person, such as facial images or dactyloscopic data.

Voice becomes biometric data under this definition only when it is processed through specific technical means for the purpose of uniquely identifying someone. That depends on what the dataset trains, not on how the audio sounds.

GDPR Article 9(1)

Processing of personal data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, and the processing of genetic data, biometric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person's sex life or sexual orientation shall be prohibited.

If your use of the voice data is biometric in the Article 4(14) sense, processing is prohibited unless an Article 9(2) condition applies. The usual one is explicit consent, which sits on top of an Article 6 basis rather than replacing it.

GDPR Article 6(1)(a)

Processing shall be lawful only if and to the extent that at least one of the following applies: (a) the data subject has given consent to the processing of his or her personal data for one or more specific purposes.

Consent is the basis most voice projects rely on, and the word doing the work is "specific". Agreeing to be recorded is not agreeing that the voice will train a model, be distributed to third parties, or leave the country.

GDPR Article 13(1)(c) and 13(2)(b)

13(1)(c) the purposes of the processing for which the personal data are intended as well as the legal basis for the processing; 13(2)(b) the existence of the right to request from the controller access to and rectification or erasure of personal data or restriction of processing.

Article 13 puts the information duty on whoever collects the data, so a buyer who was not the collector inherits the consequences of that notice being missing, and it is the most common defect in a voice dataset.

Who it applies to

GDPR reaches a buyer in two ways. If you or the seller are established in the EU, Article 3(1) applies to processing in the context of that establishment. If neither of you is, Article 3(2) can still reach you where the processing relates to offering goods or services to people in the EU or to monitoring their behavior there. The safest assumption for a dataset containing EU speakers is that GDPR questions attach to the chain, because the collector who recorded them in the EU was subject to it in any event.

In practice the constraint that binds most deals is downstream rather than regulatory. EU enterprise customers will not accept a voice corpus without a documented basis, and a US buyer selling into Europe will be asked to demonstrate the same chain by its own customers. That is true regardless of whether a supervisory authority ever looks at the file.

Two categories need separate handling rather than the same form with a different name. Children's speech requires guardian consent and a stated permitted use, and it is the category most likely to be scrutinised. Data collected from call centers or telephony is usually unusable, because the other party to the call was never asked.

  • Recordings made for the purpose of training, with speaker consent: the straightforward case, and the one the checklist below is written for.
  • Repurposed archival or broadcast audio: the original consent almost never covers AI training, so the analysis usually fails.
  • Call recordings and telephony data: one party is generally missing from the consent chain, which is why we decline these projects.

What it costs to get wrong

GDPR sets two tiers of administrative fine. Article 83(4) covers a range of obligations at up to 10 million euros or 2% of total worldwide annual turnover, whichever is higher. Article 83(5) covers the basic principles and the conditions for consent under Articles 5, 6, 7 and 9 at up to 20 million euros or 4%, whichever is higher. A voice dataset with a missing or invalid basis sits in the higher tier.

The second exposure is a compensation claim. Article 82 gives data subjects a right to compensation for material and non-material damage, and it is asserted individually or through collective actions. For a corpus with thousands of speakers, that is a per-person claim rather than a single regulatory penalty.

Enforcement on voice-adjacent processing has been concrete rather than theoretical. The French CNIL fined Clearview AI 20 million euros in 2022, the Dutch data protection authority fined it 30.5 million euros in 2024, and the Italian Garante fined OpenAI 15 million euros in December 2024. These are not voice datasets, but they establish that biometric-adjacent processing without a valid basis is an enforcement priority.

The practical cost is often contractual. An erasure request you cannot service, because the files carry no speaker identifier, turns a compliance question into a product defect that cannot be fixed retroactively.

How to comply when you are buying data

A defensible voice dataset is built as a chain, and every link is a document. The questions below are the ones that decide whether the chain holds when someone else reads it.

On projects involving EU speakers we build consent documentation to these requirements and deliver it with the dataset rather than holding it back, and the transfer mechanism is addressed in the agreement rather than left implicit. We also decline work where the chain cannot be made sound, including call recordings and any dataset where the speakers were not told what the data would be used for.

We are a sourcing company, not a law firm. Nothing here is legal advice, and whether a particular dataset is permissible depends on your use case and your jurisdiction. Our role is to make the facts of the chain accurate and legible so that your counsel can assess it.

  • Get the lawful basis in writing, and make it match the use. If the basis is consent, the consent text should name AI or machine learning training as a purpose, not merely recording or research.
  • Ask for the Article 13 notice, not just the signature page. A signed form that never told the speaker the purpose is a weak consent, and the notice is the document that shows the purpose was communicated.
  • Check the transfer mechanism if the data crosses a border. EU speakers plus non-EU training infrastructure is a restricted transfer and needs its own legal basis on top of the consent.
  • Require a speaker identifier on every file. Without it, an access or erasure request cannot be actioned, and the dataset cannot be repaired once it has been distributed.
  • Ask whether the dataset is intended for identification or verification. If it is, the Article 9 analysis applies and you need an explicit consent record, not a general one.

How we handle consent and licensing →

Related compliance topics

Not legal advice

We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.

Sourcing data under Voice Data under GDPR?

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com