Global

AI Training Data and Copyright

The short answer. Two different acts get confused here. Ingesting a work to train a model is a reproduction, argued in the US as fair use and in the EU under a text-and-data-mining exception that the rights holder can switch off. Emitting something substantially similar to a specific training item is a separate act, and it is much harder to defend. No US appellate court has settled whether training is fair use. Treat any vendor who tells you it is settled as describing a litigation position rather than a clearance, and check which of the two acts your product actually performs.

The law

17 U.S.C. §101 (definition of "derivative work")

A "derivative work" is a work based upon one or more preexisting works, such as a translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgment, condensation, or any other form in which a work may be recast, transformed, or adapted.

This definition is where the training-versus-output split starts. Model weights are not a copy of any single training item in the way a duplicated file is, which is why the ingestion step is argued as reproduction rather than as preparation of a derivative work. The derivative analysis becomes much harder to avoid on the output side, when a system produces something substantially similar to a particular item it was trained on.

17 U.S.C. §107

In determining whether the use made of a work in any particular case is a fair use the factors to be considered shall include — (1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (4) the effect of the use upon the potential market for or value of the copyrighted work.

Fair use is a defense, not a permission. It is raised after a claim exists, it is decided case by case, and the fourth factor is the one that has moved most in the recent litigation. For a buyer, the practical consequence is that a supplier claiming "training is fair use" is offering an argument you would have to run yourself, with your own corpus as the exhibit.

Directive (EU) 2019/790 Article 4(3)

The exception or limitation provided for in paragraph 1 shall apply on condition that the use of works and other subject matter referred to in that paragraph has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online.

In the EU the default leans toward permitted, but only until a rights holder reserves their rights, and the reservation can be made by machine-readable means on content published online. A dataset described as "publicly available" tells you nothing about whether this condition holds. Separately, Article 4(2) allows the copies to be retained only for as long as necessary for mining, which is a narrower right than owning a corpus.

Who it applies to

This page is about material that carries copyright — text, audiobooks, music, broadcast audio, published transcripts. It is not about the speaker's own rights in their voice, which is a separate question governed by data protection and publicity law rather than by copyright.

The territorial split is the part buyers underestimate. Fair use is a US doctrine and it is a defense, so it is asserted in a US proceeding after a claim is filed. EU law does not have fair use; it has a closed list of exceptions, and the two text-and-data-mining entries on that list are the only ones that reach training. Where the copying happens matters less than where the material is protected, so a training run on US servers does not remove the EU analysis from an EU-sourced corpus.

The output question follows your product, not your training pipeline. A model that reproduces recognizable passages, generates a voice close enough to a specific performer, or reconstructs a melody is doing something the ingestion exception does not cover. Buyers who clear the training step and ignore the output step have cleared the cheaper half.

  • Training on lawfully accessed material, then selling a model: the ingestion analysis plus the distribution question.
  • Training on material you scraped past a reservation of rights: the exception in Article 4(3) is not available, so the analysis falls back to a license you do not have.
  • Training on licensed material but shipping a system that emits near-copies: a separate exposure that the training license may not address at all.

What it costs to get wrong

The statutory exposure in the US is the same machinery as any other infringement claim: actual damages and the infringer's profits under 17 U.S.C. §504(a), or statutory damages instead under §504(c) of 750 to 30,000 dollars per work, rising to 150,000 dollars per work for willful infringement. Fees are recoverable under §505. The multiplier that matters is the work count, because a corpus is a collection of separate works.

The larger practical exposure is injunctive. A court can order a model or a dataset taken down or destroyed, and unlike a damages award that is not something you can absorb and continue past. For a product already in customers' hands, that is the outcome legal teams plan around.

In the EU the remedy is national but the direction is set: a German court in 2026 found unlicensed use of copyrighted music in AI training infringing, after the Hamburg Regional Court had found the research-purpose LAION reproduction lawful in 2024. The line between those two outcomes is purpose and license, not the technology.

Reputationally, the cost lands on the disclosure obligation. Under the EU AI Act, providers of general-purpose models must publish a summary of training content and a copyright policy, so a rights problem in the corpus becomes a public statement rather than a private dispute.

How to comply when you are buying data

Most of the risk here is settled by where you source, not by how you argue. The steps below are the ones that change the answer, and they are ordered by how much exposure they remove rather than by what they cost.

On the speech side, the practical test is whether the recordings were made for the purpose of training. Read-aloud corpora, commissioned recordings and studio sessions usually carry clean rights. Broadcast archives, scraped podcasts and user uploads usually do not, and the difference is visible in the paperwork rather than in the audio.

We are a sourcing company, not a law firm, and this page is documentation rather than legal advice. What we can do is tell you where each hour came from, what the speaker agreed to, and what rights the source carries — so that your counsel is assessing facts instead of assumptions.

  • Prefer licensed or commissioned sources over scraped ones. The cost difference is usually smaller than the cost of an after-the-fact rights review on a corpus you have already integrated.
  • Ask where the text or script came from, separately from where the audio came from. Read-speech datasets frequently have clean audio and borrowed scripts, and the borrowed half is the exposure.
  • Check for a reservation of rights on any web-derived text. Article 4(3) means the EU exception can be absent from content that looks freely usable.
  • Separate the training grant from the output grant in your contract. Ask for an explicit statement about what the model may generate, and whether the supplier warrants that no training item will be reproduced.
  • Require the supplier to disclose known claims. A vendor who has already received a takedown or a demand letter is a vendor whose corpus carries a known defect, and you want that disclosed before signature rather than after.

How we handle consent and licensing →

Related compliance topics

Not legal advice

We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.

Sourcing data under AI Training Data and Copyright?

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com