EU
EU AI Act Transparency Requirements
The short answer. Two different transparency regimes run side by side, and they are easy to confuse. Article 50 imposes disclosure duties on the providers and deployers of certain AI systems — telling people they are talking to a machine, marking synthetic output, and disclosing emotion recognition. Article 53 imposes documentation and disclosure duties on providers of general-purpose models, including a public summary of training content and a copyright policy. The dates are phased: general-purpose obligations applied from 2 August 2025 and the Article 50 duties from 2 August 2026.
The law
Regulation (EU) 2024/1689 Article 50(1)
Providers shall ensure that AI systems intended to interact directly with natural persons are designed and developed in such a way that the natural persons concerned are informed that they are interacting with an AI system, unless this is obvious from the point of view of a natural person who is reasonably well-informed, observant and circumspect, taking into account the circumstances and the context of use.
This duty lands on whoever places the system on the market, not on whoever supplied the training data. A voice assistant built on a purchased corpus owes this disclosure, and the corpus documentation is what supports the system's technical file.
Regulation (EU) 2024/1689 Article 53(1)(c)
(c) put in place a policy to comply with Union law on copyright and related rights, and in particular to identify and comply with, including through state-of-the-art technologies, a reservation of rights expressed pursuant to Article 4(3) of Directive (EU) 2019/790.
This is the copyright policy obligation for general-purpose model providers, and it points back at the text-and-data-mining opt-out. It requires an actual mechanism for identifying reservations of rights, not a statement of intent. The practical effect is that suppliers must now document where training content came from and what rights position it carried, which is a provenance requirement in all but name.
Regulation (EU) 2024/1689 Article 53(1)(d)
(d) draw up and make publicly available a sufficiently detailed summary about the content used for training of the general-purpose AI model, according to a template provided by the AI Office.
This obligation turns the contents of a training corpus into a public disclosure. It is a summary rather than a file listing, but it has to be sufficiently detailed and it follows a template issued by the AI Office. A buyer training a general-purpose model will therefore need enough information about your corpus to write it.
Regulation (EU) 2024/1689 Article 53(2)
The obligations set out in paragraph 1, points (a) and (b), shall not apply to providers of AI models that are released under a free and open-source licence. ... This exception shall not apply to general-purpose AI models with systemic risks.
The open-source exemption covers technical documentation and the information given to downstream integrators. It does not cover the copyright policy in point (c) or the training-content summary in point (d), and it disappears for models with systemic risk. An open-source release is not an escape from the two obligations that concern training data.
Who it applies to
The regulation applies to providers placing systems on the EU market wherever they are established, and to deployers established or located in the EU. A US company that puts a model on the EU market is a provider for these purposes, and a US company whose system is used inside the EU by an EU entity will often be caught through the deployer side. Establishment is not the test.
The two regimes bind different parties. Article 50 duties fall on providers for the disclosure and marking obligations, and on deployers for the emotion recognition and deepfake disclosures. Article 53 duties fall only on providers of general-purpose models. A company that fine-tunes someone else's model is generally not the provider of that model, but may be a provider of an AI system and therefore inside Article 50.
The dates are phased rather than single. The prohibitions and the AI literacy duty applied from 2 February 2025, the general-purpose model obligations in Chapter V from 2 August 2025, and the Article 50 transparency obligations from 2 August 2026, which is now in the past. The Commission published guidelines on the Article 50 duties and recognized a transparency code of practice in July 2026.
A data supplier sits outside all of this directly. The obligations are on providers and deployers, not on the people who sold them the corpus. The indirect effect is what matters commercially: buyers now need documentation they did not previously ask for.
- Provider of a general-purpose model: Article 53 duties, including the copyright policy and the training-content summary.
- Provider of an AI system built on a purchased corpus: Article 50 duties, plus the obligation to know what is in the training data.
- Data supplier: no direct obligation, but the buyer's ability to comply depends on what you can document.
What it costs to get wrong
The fine tiers are set by Article 99. Non-compliance with the Article 50 transparency obligations falls under Article 99(4)(g) at up to 15 million euros or 3% of total worldwide annual turnover, whichever is higher. The same tier covers a range of provider, importer and distributor obligations.
The upper tier is reserved for the prohibited practices in Article 5, at up to 35 million euros or 7% under Article 99(3). Supplying incorrect, incomplete or misleading information to notified bodies or authorities falls in between at up to 7.5 million euros or 1% under Article 99(5). Where the offender is an SME, Article 99(6) provides that the lower of the fixed amount and the percentage applies.
For general-purpose model providers there is a separate enforcement route. Article 101 allows the Commission to impose fines not exceeding 3% of annual total worldwide turnover or 15 million euros, whichever is higher, where a provider intentionally or negligently infringes the relevant obligations. That is a Commission-level power rather than a national one, which changes the shape of the enforcement risk.
The exposure that arrives first is usually not a fine. Buyers ask for the training-content summary inputs and the copyright policy before they will integrate a model, and a provider who cannot produce them is a provider with a blocked sale rather than a pending penalty.
How to comply when you are buying data
A data buyer's obligations under these provisions are satisfied or broken by the documentation they hold about the corpus, so the checklist below is mostly about what you ask a supplier to deliver rather than what you build.
One question decides your position, and it is worth answering before you buy anything: are you placing a model on the EU market, or deploying one that someone else placed there? The first makes you a provider with documentation duties, the second makes you a deployer with disclosure duties, and a company can be both for different products.
Where a project is destined for a general-purpose model, we can supply the source-level detail that the training-content summary requires and a written statement of the rights position on each source, and we will say plainly when a source cannot be documented. We are a sourcing company and not a law firm, so the compliance assessment for your own system remains one for your counsel.
- Establish which role you are in before you buy. Provider, deployer or neither determines which articles apply, and the answer changes what you need from the supplier.
- If you are training a general-purpose model, ask suppliers for the source inventory and the rights position per source. That is the raw material for both the Article 53(1)(c) policy and the Article 53(1)(d) summary.
- If you are building a system that talks to people, plan the Article 50(1) notice into the product, and keep the corpus documentation that supports it.
- If you are deploying emotion recognition or biometric categorisation, the Article 50(3) disclosure duty sits on you as deployer, and the notice should be reflected in what the speakers agreed to when they were recorded.
- If you generate synthetic audio, plan the Article 50(2) marking obligation into the pipeline, and check whether the license from your data supplier permits the synthetic output you intend to produce.
Related compliance topics
-
EU AI Act Article 50
eu ai act article 50
-
Data Provenance
data provenance
-
Data Licensing Agreement
data licensing agreement
-
AI Training Data Laws
ai training data laws
-
AI Training Data Governance
ai training data governance
-
California AI Training Data Transparency Act
ai training data transparency act
Not legal advice
We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.
Sourcing data under EU AI Act Transparency Requirements?
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.