How to read the US Copyright Office AI reports without over-reading them
The Office has published a series of reports on AI and copyright. Some of what they contain is examination practice that binds applicants; much of it is a recommendation to Congress. Telling the two apart is the skill.
What the reports are, and what they are not
The US Copyright Office runs the registration system and advises Congress. Its work on AI has produced registration guidance and a series of reports, published in separate parts, covering digital replicas, the copyrightability of AI-assisted output, and the training of generative models. The parts were released at different times and answer different questions, so a citation to "the Copyright Office report" without naming the part usually signals that nobody has read it.
None of these documents is a statute or a regulation. A report states the Office's analysis and, where it thinks the law is inadequate, recommends that Congress act. Courts may find it persuasive and are not bound by it, and it creates no obligation for a company buying data.
The exception is registration practice. When the Office says applicants must disclose AI-generated material in an application, that is how applications are examined, and a failure to disclose is a defect in the registration itself rather than a matter of opinion.
Signal one: does the report ask Congress to act?
The fastest way to separate a conclusion from a gap is to look for a legislative recommendation. When a report finds that current law does not reach a practice and recommends a new statute, it is telling you the practice is unregulated at federal level today.
The report on digital replicas is the clearest example. It surveyed the patchwork of state laws and concluded that a federal right is needed — which is a statement that no federal right currently exists, and that protection depends on which state the person was in.
For a data project that reading cuts in both directions. Nothing federal forbids the practice today, and nothing federal will protect a licence taken today if the law changes later. Contracts drafted in that gap should say which side bears the cost of a change in the law, because neither party can control it.
Signal two: registration guidance is narrower than it looks
Registration guidance is about what can be registered, not about what may be done with material. It answers two questions: is this output protectable, and what must the applicant disclose?
The recurring position is that copyright protects human expression, and material generated by a machine without human creative contribution does not qualify. The Office has also required applicants to disclose AI-generated content rather than list a human author over it.
Two consequences follow for a corpus:
- Synthetic material mixed into a training corpus has a thin rights story. The supplier may hold no copyright in it, which limits what the supplier can grant — a licence cannot convey more than the licensor owns.
- A supplier's claim to "own" a corpus containing generated segments should be tested against what the Office says is protectable, rather than accepted because the file exists.
- A corpus with a large synthetic share also has a disclosure problem, because the synthetic portion is a category that a training content summary has to name.
Signal three: where a report analyses facts rather than law
The report on training is the one most often misquoted, in both directions. It works through how models are trained and how the fair use factors apply, and concludes that the question cannot be answered categorically — that it depends on the source, on how the material was acquired, and on what the model does with it.
What that means in practice: the report is a map of which facts matter, not a safe harbour and not a finding of infringement. A summary claiming the Office said training is fair use, and one claiming the Office said training infringes, are both wrong. The second misreading is the more expensive one, because it leads companies to assume that no lawful path exists.
The discipline generalises to any report in the series: separate the factual findings about how the technology works, from the legal conclusion, from the policy recommendation. Those three sit in the same document and carry different weight.
How this should change a licensing file
A few habits follow from reading the reports properly rather than at second hand:
- Record the facts a court or a filing would ask about — source, acquisition method, output behaviour — at the point of purchase, because they cannot be reconstructed from the audio.
- Treat synthetic content as its own line in the rights schedule rather than blending it into a corpus description.
- For any voice or likeness element, assume a licence must name replicas explicitly, since the federal gap is exactly where state law and contract have to do the work.
- Re-read each new part as it appears instead of relying on a summary written the year the first one came out.
- Keep the primary documents in the project file. A report is short enough to read, and quoting it accurately is cheaper than being corrected by a counterparty who has.
The one-line version
A report tells you how the Office would analyse a question and what it thinks Congress should do about it. It does not tell you that a practice is lawful, and it does not bind a court. Used that way, it is a useful input to a rights file. Used as a clearance, it is a citation that will not hold.
This is an operational overview rather than legal advice, and the reports themselves are short enough to read directly. For a specific project, the analysis belongs with counsel working from the current text.