EU
EU AI Act Article 50
The short answer. Article 50 is the transparency article, and it contains four separate duties. Providers must tell people they are interacting with an AI system, and must mark synthetic audio, image, video and text output as machine-readable and detectable. Deployers must inform people exposed to emotion recognition or biometric categorisation, and must disclose deepfakes and AI-generated text on matters of public interest. Each duty has its own exceptions, and the disclosure has to come at the latest at first interaction or exposure. It applied from 2 August 2026.
The law
Regulation (EU) 2024/1689 Article 50(1)
Providers shall ensure that AI systems intended to interact directly with natural persons are designed and developed in such a way that the natural persons concerned are informed that they are interacting with an AI system, unless this is obvious from the point of view of a natural person who is reasonably well-informed, observant and circumspect, taking into account the circumstances and the context of use.
The obligation is on the provider and it attaches at design time, not at deployment. The exception is an objective test about what a reasonable person would assume, which for a voice assistant is a high bar: people routinely treat fluent synthetic speech as human.
Regulation (EU) 2024/1689 Article 50(2)
Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated.
Audio is named first, and the requirement is a machine-readable mark rather than a spoken disclaimer. The article requires the technical solutions to be effective, interoperable, robust and reliable as far as technically feasible, with exceptions for assistive editing that does not substantially alter the input. A synthetic speech product needs marking built into the pipeline, not added later.
Regulation (EU) 2024/1689 Article 50(3)
Deployers of an emotion recognition system or a biometric categorisation system shall inform the natural persons exposed thereto of the operation of the system, and shall process the personal data in accordance with Regulations (EU) 2016/679 and (EU) 2018/1725 and Directive (EU) 2016/680, as applicable.
This one sits on the deployer, which is usually the customer rather than the vendor, and it is directly relevant to voice emotion datasets. Two duties arrive together: notice to the people exposed to the system, and compliance with data protection law. Article 50(3) does not replace the GDPR analysis, it sits on top of it.
Regulation (EU) 2024/1689 Article 50(4)
Deployers of an AI system that generates or manipulates image, audio or video content constituting a deep fake shall disclose that the content has been artificially generated or manipulated.
The deepfake disclosure sits on deployers and is separate from the marking duty on providers, so the same synthetic audio can trigger both. The article carves out evidently artistic, satirical or fictional works, limiting the duty to a disclosure that does not hamper enjoyment of the work, and imposes a separate duty to disclose AI-generated text on matters of public interest.
Who it applies to
The article splits by role. Paragraphs 1 and 2 bind providers, which means whoever places the system on the EU market, wherever they are established. Paragraphs 3 and 4 bind deployers, which means the organization using the system in the EU. A company can be both at once for different products, and the two sets of duties are not interchangeable.
Timing is set by Article 50(5), which requires the information to be provided in a clear and distinguishable manner at the latest at the time of the first interaction or exposure, and requires it to meet applicable accessibility requirements. A disclosure buried in terms of service or shown after the user has already engaged does not meet that standard.
Territorially, the article follows the market rather than the company. A US provider placing a synthetic voice product on the EU market is inside paragraph 2. A US company deploying emotion recognition on EU users is inside paragraph 3. There is no establishment requirement in either case.
Two practical points for anyone buying voice data. First, paragraph 2 means synthetic output has to be markable, which is a pipeline requirement that is easier to design in than to retrofit. Second, paragraph 3 means the people whose voices are used to train an emotion recognition system may be entitled to be informed that such a system is operating, which is a purpose the consent documentation should have covered at collection time.
- Synthetic speech product placed on the EU market: provider duties under 50(1) and 50(2).
- Emotion recognition deployed on EU users: deployer duties under 50(3), plus the data protection analysis.
- Deepfake or public-interest AI text published in the EU: deployer duties under 50(4).
What it costs to get wrong
Article 50 non-compliance is enforced under Article 99(4)(g) at up to 15 million euros or 3% of total worldwide annual turnover, whichever is higher. That is the middle tier, above the informational tier and below the prohibited-practices tier of 35 million euros or 7% under Article 99(3).
Enforcement became live on 2 August 2026, when the Article 50 obligations took effect. The Commission published final guidelines on the transparency obligations and recognized a transparency code of practice as adequate in July 2026, so the interpretive position is more settled than it was when the article was drafted.
The reputational exposure is specific to this article, because a transparency failure is visible to the public rather than only to a regulator. A synthetic voice that is not marked, or a chatbot that does not disclose itself, becomes a press story before it becomes an enforcement file, and the story is what customers react to.
For a data supplier the consequence is indirect but concrete. A buyer building a product that must mark its output will look for a data license that permits the marking and the synthetic output, and a license that is silent on synthetic use stalls the deal rather than creating a fine.
How to comply when you are buying data
The article is short, and the work it creates is mostly product design plus one item of paperwork that data buyers tend to miss: the recording consent has to anticipate the disclosure duties that the finished system will carry.
The two provider duties and the two deployer duties are often held by the same company at different points in the supply chain, and the disclosure wording has to work in both places. A supplier that ships synthetic audio to a customer who then publishes it is a provider, and the customer is a deployer. Both have obligations on the same file.
On voice emotion and identification projects we record the intended purpose in the consent documentation, including the fact that the resulting system may disclose its operation to people it is used on, and we flag projects where the intended purpose cannot be reconciled with what the speakers were told. We are a sourcing company rather than a law firm, so the compliance assessment of your finished system is one for your counsel.
- Decide which paragraphs apply to you before you buy data, because the answer changes what you need in the license and in the consent record.
- If you generate synthetic speech, check that the license expressly permits the creation and distribution of synthetic output, and plan the machine-readable marking into the pipeline.
- If you train emotion recognition or biometric categorisation, make sure the speakers were told that the system may be disclosed to the people it is used on. That is a purpose, and it belongs in the consent.
- Write the timing of your disclosure into the product specification. The article requires it at first interaction or exposure, which is a product decision rather than a legal one.
- Keep the corpus documentation that supports the system's technical file. The disclosure duties are on you, but the evidence that you understood your training data is what a reviewer will ask for.
Related compliance topics
-
Data Provenance
data provenance
-
Data Licensing Agreement
data licensing agreement
-
AI Training Data Laws
ai training data laws
-
AI Training Data Governance
ai training data governance
-
California AI Training Data Transparency Act
ai training data transparency act
-
Biometric Data under GDPR
gdpr biometric data
Not legal advice
We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.
Sourcing data under EU AI Act Article 50?
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.