Fair use and AI training data: the argument, the open questions, and what a buyer can control
Fair use is a defence argued after the fact, factor by factor. Here is how the four factors map onto training, why no appellate court has settled the question, and where a buyer's leverage actually sits.
Fair use is a defence, not a permission
Fair use is a US doctrine that excuses a use of copyrighted material which would otherwise infringe. It is raised in response to a claim, decided on the facts of one case, and it produces no certificate, no clearance and no advance ruling. There is no office that tells you in advance that your training run is covered.
That structure explains the most common mistake in this area: treating fair use as a property of a use rather than an outcome of a dispute. Training a model on copyrighted material is either fair use or it is not, and the answer is produced by a court, after a claim, if a claim is ever brought.
It also explains what a supplier's assurance is worth. A vendor who says training is fair use is describing a litigation position. What the buyer receives is the position, not the outcome, and the corpus is the exhibit.
The four factors, applied to training
The statute lists four considerations and assigns them no weights. Each one behaves differently when the use is model training:
- Purpose and character of the use. Training has been argued as transformative, because a model learns statistical patterns rather than republishing the work. The commercial nature of the use pushes the other way, and no court has treated "the output is different" as decisive on its own.
- Nature of the work. Published, factual material fares better than unpublished or highly creative material. For a corpus built from fiction, scripts or performances, this factor is unfavourable and there is little to argue about.
- Amount used. Training ingests complete works, which looks bad in isolation. The counter-argument is that no single work is expressed in the weights — that what was taken were patterns rather than expression. Courts have engaged with that argument seriously, which is not the same as accepting it.
- Effect on the market. The contested factor, and the one that has moved most. Rights holders argue that a licensing market for training material exists, or would exist, and that training displaces it. That argument has been persuasive in some decisions and rejected in others.
Why it is not settled
No appellate court has ruled on whether training on copyrighted works is fair use. Nothing in this area binds beyond the district that decided it, and district courts have reached opposite conclusions on comparable technologies.
The outcomes have turned on facts that have little to do with the technology itself:
- How the material was acquired. Lawfully purchased or licensed copies have fared far better than material taken from shadow libraries or scraped past a technical restriction. Acquisition is where the cases have been lost.
- What the model does at output time. A system that can reproduce recognisable passages is in a different position from one that cannot, and the difference is a design decision as much as a legal one.
- The type of work. Books, news content and legal research material have produced different results, because the fourth factor is about the market for that kind of material.
The speech-data wrinkle
Audio corpora have features that change the analysis, and they are usually missing from discussions built around text:
- Read-aloud material carries two rights at once — the recording and the underlying text. Clearing the studio does not clear the script, and the script is the half that gets forgotten.
- A recording of a performance may be governed by a contract: a union agreement, a studio release, a licence from the original production. Fair use does not override a contract, so a corpus can be defensible on copyright grounds and still be unusable.
- The performer's interest in their own voice is not a copyright question. It sits in publicity and data protection law, where the fair use factors are not the test at all.
Where the buyer's leverage actually is
A buyer cannot make the fair use question disappear. A buyer can decide which risks to hold, and with what instruments:
- Source selection. Commissioned and licensed material removes the question instead of arguing it, and the price difference is usually smaller than the cost of a later rights review.
- A warranty about acquisition, which is a narrower and more provable statement than a warranty about fair use. Ask how the material was obtained, not whether training is lawful.
- An indemnity that names the risk it covers. A general indemnity is usually capped and carved back until it covers very little.
- A replacement or retraining right, so a corpus with a discovered defect can be swapped rather than litigated.
- Written disclosure of known claims, repeated at each delivery rather than given once at the start.
The posture that ages well
The one thing not to do is build a compliance position that requires a court to agree with you. Contingency planning costs a fraction of the contingency, and the planning is mostly documentary: know where each hour came from, keep the acquisition records, and write down what was known at the time of purchase.
It is also worth separating the two acts. Ingestion is the contested question; reproducing protected expression in output is the weaker defence; and acquiring material improperly is the weakest position of all. A project can improve its position on the second and third without resolving the first.
This is background for procurement discussions rather than legal advice. Whether a particular use qualifies is a question for counsel, and ultimately for a court, on facts that a data supplier may or may not be able to evidence.