Global
AI Training Data Governance
The short answer. A governance process for training data is four controls, not a policy document: a lineage record that ties every delivered file to a named source and a consent artifact; an approval gate that data has to pass before it enters a training run; a retention rule that says when the data is destroyed; and a procedure for acting on an erasure request after the model has already trained. GDPR Article 5(1)(e) supplies the retention rule. EU AI Act Article 10 supplies the quality and provenance standard. Everything else is implementation.
The law
GDPR Article 5(1)(e)
Personal data shall be kept in a form which permits identification of data subjects for no longer than is necessary for the purposes for which the personal data are processed; personal data may be stored for longer periods insofar as the personal data will be processed solely for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) subject to implementation of the appropriate technical and organisational measures required by this Regulation in order to safeguard the rights and freedoms of the data subject ('storage limitation').
The dataset is the thing being stored, not just the model. A speech corpus kept after the training run that needed it is personal data held without a purpose. The research exception is narrower than it is usually presented: it requires a basis in Article 89(1) and the safeguards that go with it, and it does not cover a corpus held for a commercial product roadmap.
EU AI Act Article 10(1)
High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis of training, validation and testing data sets that meet the quality criteria referred to in paragraphs 2 to 5.
This binds the provider of a high-risk system, not the data vendor. It is still the reason buyers now ask for dataset-level documentation, because Article 10(2) requires the provider to record the data collection processes, the origin of the data, and in the case of personal data the original purpose of the collection. That section cannot be written from the buyer's own records. It has to come from whoever collected the recordings, which makes the vendor's paperwork part of your compliance file.
GDPR Article 17(1)
The data subject shall have the right to obtain from the controller the erasure of personal data concerning him or her without undue delay and the controller shall have the obligation to erase personal data without undue delay where one of the following grounds applies: the personal data are no longer necessary in relation to the purposes for which they were collected or otherwise processed; the data subject withdraws consent on which the processing is based according to point (a) of Article 6(1), or point (a) of Article 9(2), and where there is no other legal ground for the processing.
You can delete a file. You cannot delete a trained weight. What governance has to do is make the request answerable: which recordings, from which speaker, went into which training run, and what the plan is when the honest answer is "all of them". A governance process that has never once traced a speaker identifier back to the runs that consumed their audio has not been tested.
Who it applies to
Governance is not one law, and the two provisions above bind different parties. EU AI Act Article 10 attaches to providers of high-risk AI systems placed on the EU market or put into service in the EU. GDPR Article 5 attaches to any controller processing the personal data of people in the EU, regardless of where the controller is established — Article 3(2) reaches a company with no EU presence at all when the processing relates to offering goods or services to people in the Union.
For a US buyer, that means the retention rule and the erasure right apply to a corpus of EU speakers even if the training happens in Texas and the model is never sold in Europe. A corpus recorded outside the EU, from speakers outside the EU, used outside the EU, sits outside both instruments, and governance there is a contract question: what the vendor warranted and what you can prove.
The AI Act's high-risk classification is the gate for Article 10. A general-purpose speech model trained for its own sake is not a high-risk system, but the corpus used to build it is often reused later in one — a biometric identification system, an emotion recognition component, a system used in hiring. That reuse is the practical reason to write the governance record to the stricter standard from the first delivery rather than retrofitting it.
What it costs to get wrong
GDPR Article 83 sorts infringements into two tiers. Breaches of the basic principles in Article 5, which include storage limitation, sit in the higher tier under Article 83(5): administrative fines up to 20 million euros, or in the case of an undertaking, up to 4% of the total worldwide annual turnover of the preceding financial year, whichever is higher. Article 83(4) covers the obligations in Articles 25 to 39 — which includes the processor contract in Article 28 and the impact assessment in Article 35 — at up to 10 million euros or 2%, whichever is higher.
EU AI Act Article 99(4) sets fines of up to 15 million euros or 3% of total worldwide annual turnover for non-compliance with provider obligations, which is the tier that reaches the Article 10 requirements. The Article 5 prohibitions are higher, at up to 35 million euros or 7% under Article 99(3).
The larger cost is usually not a fine. A corpus without lineage cannot be defended in a diligence review, cannot be resold, and cannot be embedded in a product that a regulated customer will buy. The practical consequence of no governance is that the data is single-use, and the price paid for it is written off at the first serious question.
How to comply when you are buying data
The process below is the buyer's half. The vendor's half is the paperwork that feeds it, and the two have to be specified in the purchase contract rather than requested after delivery.
- Lineage record, per delivery: source file name, speaker identifier, recording date, country of recording, the consent artifact that covers it, and the contract under which it was supplied. Ask for this for a random sample of twenty files. If it cannot be produced for twenty, it cannot be produced for twenty thousand.
- Approval gate: a named person who signs off before a dataset enters a training run, with the dataset version recorded. A gate with a name attached is a control. A gate described in a policy is a sentence.
- Retention schedule: for each dataset, the purpose it was collected for, the date that purpose ends, and the destruction action. Raw audio, transcripts and speaker manifests often carry different clocks and should be scheduled separately.
- Erasure procedure: how you identify a speaker's records on receipt of a request, what you delete, what you cannot delete, and what you tell the requester. Run it once on a test identifier before a real request arrives.
- Contract terms: a notification duty when a speaker withdraws consent, a defined response period, an obligation to keep consent artifacts for the term plus the retention period, and cooperation with erasure requests at no additional charge.
Related compliance topics
-
California AI Training Data Transparency Act
ai training data transparency act
-
Biometric Data under GDPR
gdpr biometric data
-
Illinois BIPA
bipa
-
Data Protection Impact Assessment
dpia
-
Data Processing Agreement
data processing agreement
-
Vendor Due Diligence Checklist
vendor due diligence checklist
Not legal advice
We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.
Sourcing data under AI Training Data Governance?
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.