The Generative AI Copyright Disclosure Act: what it would require, and why it matters before it passes
The federal bill would turn the training corpus into a filing with the Copyright Office. It is not law yet, but the facts it names are the same ones buyers already ask suppliers to produce.
What the bill proposes
The Generative AI Copyright Disclosure Act is a federal proposal introduced in the US House of Representatives. It would require anyone who creates, or materially alters, a generative AI system trained on copyrighted works to file a notice with the Register of Copyrights describing the material used in training.
The content of the notice, as proposed, is a description of the copyrighted works: the categories they fall into, the sources they came from, and how they were obtained. For a system not yet released, the filing would come before it is made available to the public; for one already out, within a short window after the requirement takes effect. The bill also contemplates keeping records and civil penalties for failing to file.
Note the shape of the thing. This is a filing with a government office, not a public webpage. It is closer in kind to a registration than to the published training summaries that other regimes require, and the difference decides what a supplier can and cannot keep confidential.
It is a proposal, not a rule — and that changes how to use it
As of this writing no version of the bill has been enacted into law. A bill binds no one until it passes both chambers and is signed, and its text can change at every stage, so its provisions describe what a group of legislators wants the law to become rather than what the law is. Check the current status in the official congressional record rather than trusting a summary — including this one.
That does not make it irrelevant. Three reasons to track it anyway:
- It names the exact facts a disclosure regime needs, which is a specification for what to start capturing now.
- The same disclosure idea has already been enacted at state level in the US and at EU level, so the direction of travel does not depend on this bill passing.
- A buyer's procurement team will ask the question whether or not a statute compels it, because the answer reduces the buyer's own exposure.
How the disclosure regimes differ
Three regimes are in play and they are not copies of one another. The EU route requires a general-purpose model provider to publish a summary of training content, following an official template, addressed to the public. The US state route, already in force in one state, requires a developer to publish a training data summary on its own site, with its own list of contents. The federal proposal would add a confidential filing with the Copyright Office, describing copyrighted works used.
The differences matter to a supplier in a way that is easy to overlook: a confidential filing may need to name material that a public summary deliberately describes only by category. A contract that forbids the supplier from disclosing a source in public may still leave the buyer obliged to name it in a filing, and contracts drafted before these regimes existed rarely say which of the two applies.
The practical resolution is to write the contract so it distinguishes them: what may be published, what may be filed, and what may be shared with a regulator under confidentiality.
What a corpus owner should be able to produce
Whatever the regime, the same underlying facts get requested, and every one of them is captured at collection time or not at all:
- What categories of works are in the corpus, and roughly in what proportion to each other.
- Where each category came from — commissioned, licensed, scraped, user-contributed, synthetic.
- The rights basis for each category, and any reservation of rights attached to the source.
- Acquisition dates, so the description can be tied to a particular training run rather than to the corpus as it looks today.
- What has been removed, and why — exclusions are part of the description, not an absence from it.
Two tests worth running this week
First test: could you produce that list for a corpus delivered two years ago without interviewing anyone? If the answer depends on a person rather than a record, the record does not exist.
Second test: does the list survive contact with the files? Do the counts reconcile, do the acquisition dates match the storage timestamps, and does the rights basis for each category match the paperwork? A description that disagrees with the corpus is worse than no description, because it is the version someone will rely on.
One corpus type deserves special attention: material acquired from a broker rather than collected directly. The seller of that material may not hold the facts either, which means the disclosure gap is inherited at purchase. Discovering that before signature changes the price; discovering it afterwards changes the relationship.
The procurement consequence
Disclosure obligations that bind the buyer become questions the buyer asks the supplier. Expect them in this form: a warranty that the corpus contains no material the supplier was not entitled to use, a schedule describing the categories and sources, and a duty to notify if any of it turns out to be wrong.
Two clauses deserve attention when they are offered. A representation that survives closing, since a disclosure problem surfaces years after the deal. And a cooperation clause obliging the supplier to help answer a regulator or complete a filing, which is worth nothing if the supplier archived nothing.
This is an operational summary of a proposal and the rules around it, not legal advice. Confirm the current status of any bill, and its application to a specific system, with counsel before relying on it in a contract or a filing.