Global

AI Training Data Laws

The short answer. There is no single global rule on training data, and any page that offers one is describing a jurisdiction rather than the law. Four regimes interact: the EU AI Act, which imposes documentation and transparency duties on model providers; the GDPR, which governs the people in the data; US copyright law, where the training question is litigated under fair use rather than settled by statute; and a growing set of US state laws that mostly require disclosure rather than permission. Your obligation depends on where you operate, where you sell, and where the speakers were.

The law

GDPR Article 3(2)

This Regulation applies to the processing of personal data of data subjects who are in the Union by a controller or processor not established in the Union, where the processing activities are related to: (a) the offering of goods or services, irrespective of whether a payment of the data subject is required, to such data subjects in the Union; or (b) the monitoring of their behaviour as far as their behaviour takes place within the Union.

This is the provision that makes the speakers' location the first question in any dataset review. A company with no EU presence can still be inside the GDPR if its processing relates to offering services to people in the EU. Where the corpus is recordings of EU speakers, the collector was subject to GDPR in any case, and the buyer inherits the consequences of how that was handled.

Regulation (EU) 2024/1689 Article 2(1)

This Regulation applies to: (a) providers placing on the market or putting into service AI systems in the Union, irrespective of whether those providers are established within the Union or in a third country; (b) deployers of AI systems that have their place of establishment or are located within the Union.

The AI Act follows the market rather than the company, which is why a US provider selling into the EU is inside it. The obligations that touch training data are documentation obligations rather than permissions: the Act tells you what to disclose about your corpus, not whether you were allowed to use it.

17 U.S.C. §107

In determining whether the use made of a work in any particular case is a fair use the factors to be considered shall include — (1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (4) the effect of the use upon the potential market for or value of the copyrighted work.

The US has no statute that addresses training data. It has a defense that is argued after a claim exists, decided case by case, and still unsettled at the appellate level. Training on lawfully obtained material is widely practiced and contested, and not something a buyer can rely on as a clearance.

Who it applies to

The practical approach is to answer four questions in order. Where were the people in the data? Where does the training happen? Where will you sell the output? And what does the output do?

The first question determines the data protection regime: EU speakers bring the GDPR into the chain regardless of where you are. The second determines the copyright analysis, because copying happens where the copies are made and rights are territorial. The third determines which AI-specific regime applies, since the EU AI Act follows the market and US state disclosure laws follow the consumer. The fourth determines whether a use-specific regime attaches, such as the biometric statutes.

The US state picture is moving faster than the federal one and is mostly about disclosure. California's Generative AI Training Data Transparency Act, codified at Civil Code §3111, requires developers of publicly available generative AI systems to post a summary of their training data. These provisions generally require you to describe what you trained on rather than to obtain permission for it, which makes documentation the operative compliance work.

Two things are worth holding on to. These regimes stack rather than conflict: a dataset can satisfy the EU AI Act and fail the GDPR, or clear US copyright and fail an EU text-and-data-mining analysis. And the direction of travel everywhere is toward more documentation, which makes a provenance record built at collection time the cheapest compliance investment available.

  • EU speakers, US training, global sales: GDPR for the people, US law for the copies, the AI Act for anything sold into the EU.
  • US speakers, US training, US sales: copyright and fair use, state biometric and publicity statutes, and the California disclosure duty.
  • Anywhere to the EU market: the AI Act applies on the provider and deployer side regardless of establishment.

What it costs to get wrong

The fine figures differ by regime and by an order of magnitude, which is why the "what does it cost" question has no single answer. GDPR tops out at 20 million euros or 4% of total worldwide annual turnover under Article 83(5), and 10 million euros or 2% under Article 83(4). The EU AI Act tops out at 35 million euros or 7% for prohibited practices, 15 million euros or 3% for transparency failures, and applies separate Commission-level fines to general-purpose model providers under Article 101.

The US has no equivalent administrative tier. Exposure is civil: statutory damages of 750 to 30,000 dollars per work under 17 U.S.C. §504(c), rising to 150,000 dollars for willful infringement, plus fees; 1,000 to 5,000 dollars per violation under BIPA; and the greater of 750 dollars or actual damages under California's publicity statute. These figures look smaller individually and larger in aggregate, because they multiply by the number of works or the number of people rather than by turnover.

The largest US outcomes on record came from state attorneys general rather than regulators: the Texas settlements with Meta at 1.4 billion dollars and Google at 1.375 billion dollars were reached under a statute with no private right of action and no damages clause.

Across all of these regimes the pattern is the same. The penalty is rarely the first cost. The first cost is a deal that cannot close, a model that cannot ship, or a disclosure obligation that exposes a gap the company had not had to describe in public.

How to comply when you are buying data

Because the rules stack rather than replace each other, the efficient approach is to build one documentation set that satisfies the strictest regime you touch, rather than three sets that each satisfy one. The strictest regime is usually the one attached to the speakers, not the one attached to your product, because you can change where you sell more easily than you can change where your data came from.

On the speech side specifically, three facts about a corpus determine most of the analysis: where the speakers were, what they were told, and what rights the source material carried. We document those three on every project and can produce them on request, and we decline work where any of them cannot be established. We are a sourcing company and not a law firm, so this is a factual record rather than legal advice.

  • Map your position before you buy. Write down where the speakers are, where you train, where you sell, and what the output does. Every regime in this set keys off one of those four answers.
  • Build one provenance record and one consent chain rather than a version per market. A record built to the EU standard is the record a US enterprise buyer will ask for anyway.
  • Ask suppliers for the speaker location field. It is the single most useful fact in the file and the one most often missing from a dataset specification.
  • Check the disclosure obligations that attach to your product, not only to your data. The California training data summary and the EU training content summary are both triggered by what you build, not by what you bought.
  • Treat unsettled law as a risk to price rather than a question to answer. Where the US position on training is genuinely unresolved, the mitigation is licensed sources and documented chains, not a stronger argument.

How we handle consent and licensing →

Related compliance topics

Not legal advice

We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.

Sourcing data under AI Training Data Laws?

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com