Article 53 of the EU AI Act: what general-purpose model providers owe on training data

Article 53 puts four duties on providers of general-purpose models. Two of them reach straight into the training corpus, and both are satisfied or broken by facts a supplier holds.

Two of the four duties are about data

Article 53 is the provision of the EU AI Act that governs general-purpose AI models — the broadly capable models that other companies then fine-tune, embed or build products on top of. It places four obligations on whoever places such a model on the market.

Two of them are about the model itself: technical documentation of the training and testing process, and information handed to the companies downstream who integrate the model. Two others reach into the corpus — a copyright policy, and a public summary of the content used for training.

The split matters commercially because the two halves have different escape routes and different dependencies. The model-side duties can be discharged by the provider alone. The data-side duties cannot, because the provider can only describe a corpus using facts that were captured while it was being collected, usually by somebody else.

The training content summary, and what sufficiently detailed means

The obligation is to draw up and make publicly available a summary of the content used for training, following a template issued by the AI Office. It is deliberately a summary rather than a file manifest, and the phrase carrying all the weight is "sufficiently detailed."

Read the purpose behind the template: a reader is trying to judge what the model was exposed to and what that implies, not to rebuild the corpus. What the summary has to convey, in substance: which modalities are covered, what the main categories of data are and roughly how they weigh against each other, which languages appear, and what kinds of source the material came from — commissioned, licensed, public, user-contributed, synthetic.

The operational consequences follow from that:

  • The summary cannot be reverse-engineered from the weights. Source facts have to be recorded per batch, at collection time.
  • "Publicly available web data" is a source category, not a description. The summary still has to say what kinds of content, at what scale.
  • A source that cannot be characterised — an undocumented broker corpus, a batch with no provenance — shows up as a hole in a public document rather than as an internal gap.

The copyright policy is a mechanism, not a sentence

The second data-facing duty is to put in place a policy for complying with EU copyright law, and specifically for identifying and complying with reservations of rights expressed under the text-and-data-mining exception in the digital single market directive — the machine-readable opt-outs that rights holders attach to content published online.

The word doing the work is "identify." A policy stating that the company respects copyright is a statement, not a policy. What an examiner would look for is a described process: how rights reservations are checked at crawl or acquisition time, what technology performs the check, what was excluded as a result, how a rights holder can reach the provider, and how the policy is updated when a new source enters the corpus.

One consequence is easy to miss. A model is trained once and then reused for years, so the policy has to be able to describe the corpus as it stood at a specific training run. A policy that only describes the current pipeline cannot answer a question about a model released two versions ago.

Systemic risk raises the stakes, and open source exempts less than people assume

Above a capability threshold, a general-purpose model is treated as carrying systemic risk, and its provider takes on further duties around evaluation, adversarial testing and incident reporting. From a data standpoint the effect is that the summary and the policy are read by more people, with more adversarial intent, and with a model that is harder to modify in response.

The open-source carve-out is narrower than it is usually described. It lifts the technical documentation and downstream information duties for models released under a free and open-source licence. It does not lift the copyright policy or the training content summary, and it does not survive a systemic risk classification.

A misconception worth naming plainly: fine-tuning somebody else's model does not make you the provider of that model. But the obligations attach to whoever places a general-purpose model on the EU market, and a substantially modified model released under your own name can cross that line — which is a question for counsel, not for the data team.

What a supplier has to be able to hand over

The provider's compliance depends on facts that the data supplier holds and the provider cannot invent. The handover, per source, is short:

  • The source category — commissioned, licensed, public corpus, user-contributed — and the date the material was acquired.
  • Coverage: modality, languages, and scale in bands rather than exact counts where the contract requires it.
  • The rights basis for each category, and whether any rights holder has attached a reservation of rights to it.
  • Whether synthetic material is present, and roughly what share of the corpus it represents.
  • Anything that genuinely cannot be disclosed, stated as a stated limit rather than left blank.

Where this lands in a purchase

The practical test for a buyer is whether the supplier could support the disclosure today, not whether it promises to help later. A supplier that answers a source-inventory request in a week, from records, is in a different category from one that answers it in a month, from memory.

The clauses that carry the obligation are worth reading twice: a cooperation clause that obliges the supplier to help answer a regulator or complete a filing, a notification duty if a source turns out to be something other than what was described, and an allocation of who bears the cost if a corpus has to be replaced. A disclosure obligation is not a legal question the buyer can solve alone, because the facts live with the seller.

This is an operational reading of the provision for data teams rather than legal advice. Whether a particular model is in scope, and what its summary must contain, are decisions for counsel working from the current text and any guidance issued since.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com