Global

Consent for AI Training

The short answer. Consent taken for one purpose does not automatically cover training a model. GDPR Article 5(1)(b) requires data to be collected for specified, explicit and legitimate purposes and not further processed incompatibly, and Article 4(11) requires consent to be specific and informed. Where the original basis was consent, the practical test is whether the speaker, at the moment of recording, would have expected their voice to end up in a model that may be distributed. If the form did not say so, the answer is usually no.

The law

GDPR Article 5(1)(b)

collected for specified, explicit and legitimate purposes and not further processed in a manner that is incompatible with those purposes; further processing for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes shall, in accordance with Article 89(1), not be considered to be incompatible with the initial purposes ('purpose limitation').

Purpose limitation is the principle that decides most AI training consent questions. The purpose has to be specified and explicit at collection. A release signed for "voice data collection" specifies nothing about training, distribution or sublicensing, and each of those is a further processing operation that has to be tested.

GDPR Article 4(11)

'consent' of the data subject means any freely given, specific, informed and unambiguous indication of the data subject's wishes by which he or she, by a statement or by a clear affirmative action, signifies agreement to the processing of personal data relating to him or her.

Specific and informed are doing the work here. A form that says data may be used for research and development does not name model training, does not name distribution, and does not name sublicensing. Whether it covers any of them is a question about what a reasonable speaker would have understood, and the usual answer is that it covers none of them clearly.

GDPR Article 6(4)

Where the processing for a purpose other than that for which the personal data have been collected is not based on the data subject's consent or on a Union or Member State law which constitutes a necessary and proportionate measure in a democratic society to safeguard the objectives referred to in Article 23(1), the controller shall, in order to ascertain whether processing for another purpose is compatible with the purpose for which the personal data are initially collected, take into account, inter alia: any link between the purposes for which the personal data have been collected and the purposes of the intended further processing; the context in which the personal data have been collected, in particular regarding the relationship between data subjects and the controller; the nature of the personal data.

This is the compatibility route for a new purpose, and it applies where the further processing is not based on consent or on law. Two of the listed factors usually decide the answer for a speech corpus. The context of collection — a speaker recruited for a paid session has different expectations from a user who uploaded a sample to a free tool. And the consequences, because a voice in a distributed model cannot be recalled.

Who it applies to

The question follows the speakers, not the vendor. A corpus recorded in the EU and delivered to a US buyer is within the scope of the GDPR, and Article 3(2) reaches a controller outside the EU whose processing relates to offering goods or services to people in the Union, which a model distributed into the EU is. The compatibility analysis has the same shape whether the training happens in one country or three.

One point is worth being precise about, because it is widely misunderstood. Article 6(4) does not rescue a consent-based collection. Where the original processing was lawful because the speaker consented, a new purpose that the consent does not cover needs a basis of its own, and in practice that means going back to the speaker or finding a different basis entirely. A vendor that says the original consent covers model training should be asked to show the sentence in the form.

The reverse case is also common and also worth checking. Where the original collection rested on a legitimate interest or on law rather than consent, the compatibility test in Article 6(4) is the right instrument, and the answer can be yes. That is a real route, but it has to be argued with the factors in the provision rather than assumed from the fact that the data was lawfully collected.

What it costs to get wrong

Training on data without a valid basis puts the processing under Articles 5, 6 and 7, which sit in the higher tier under Article 83(5): administrative fines up to 20 million euros, or in the case of an undertaking, up to 4% of the total worldwide annual turnover of the preceding financial year, whichever is higher.

The operational problem is worse than the fine. A trained model cannot be un-trained. The remedy for a missing basis is to retrain without the data or to stop distributing the model, and both costs land on the buyer rather than on the vendor that supplied the corpus. That asymmetry is why the consent wording is worth reading before the purchase order rather than after the first training run.

The contractual route is a representation about the scope of the consents, an indemnity, and a notification duty if a speaker withdraws. The diligence file — the form as it was presented, the explanation given, the withdrawal records — is the evidence when the question is tested.

How to comply when you are buying data

Consent scope is a document question, and it is answerable in an afternoon if the vendor has the documents.

  • Read the consent form before the delivery, not after. Look for four words: training, distribution, sublicensing, derivative works. If a form does not contain them, assume the consent does not cover them.
  • Ask for the form as it was actually presented, in the language the speaker read. A translated summary prepared for the buyer is not the form the speaker signed.
  • Establish what the original basis was. If it was consent, do not accept a compatibility argument as a substitute for consent that names training, because consent has to be specific to the purpose it covers.
  • Ask the vendor to state in the contract whether the consents cover training, distribution and sublicensing, separately, and to identify which of them they do not cover. A specific no is more useful than a general yes.
  • Check the withdrawal mechanism. Consent can be withdrawn at any time, and Article 17(1)(b) makes withdrawal a ground for erasure, so the contract needs a notification duty and a defined response period.
  • If the consents do not cover the intended use, price the alternative before committing: re-recording with a proper form, or a licensing arrangement with the right to train. Both are cheaper than a model that cannot be distributed.

How we handle consent and licensing →

Related compliance topics

Not legal advice

We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.

Sourcing data under Consent for AI Training?

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com