Global
AI Training Data Licensing
The short answer. Training a model on someone else's work means making copies of it, and the reproduction right belongs to the rights holder. The EU has two text-and-data-mining exceptions, but one is limited to research organizations and the other can be switched off by the rights holder. The US has no training-specific exception at all. A license is how you replace that uncertainty with a document. The three things to check are which rights the grant actually names, whether it covers the compiled corpus as well as its contents, and whether it survives into the models you ship.
The law
17 U.S.C. §106(1)
Subject to sections 107 through 122, the owner of copyright under this title has the exclusive rights to do and to authorize any of the following: (1) to reproduce the copyrighted work in copies or phonorecords.
Building a training set from protected works means reproducing them — onto a disk, into a corpus, into the copy your vendor delivers. That reproduction is the rights holder's exclusive right, so the question is whether you hold a permission or fit an exception.
Directive (EU) 2019/790 Article 3(1)
Member States shall provide for an exception to the rights provided for in Article 5(a) and Article 7(1) of Directive 96/9/EC, Article 2 of Directive 2001/29/EC, and Article 15(1) of this Directive for reproductions and extractions made by research organisations and cultural heritage institutions in order to carry out, for the purposes of scientific research, text and data mining of works or other subject matter to which they have lawful access.
This is the exception most often quoted when someone argues that training is covered. It is confined to research organizations and cultural heritage institutions acting for scientific research.
Directive (EU) 2019/790 Article 4(1)
Member States shall provide for an exception or limitation to the rights provided for in Article 5(a) and Article 7(1) of Directive 96/9/EC, Article 2 of Directive 2001/29/EC, Article 4(1)(a) and (b) of Directive 2009/24/EC and Article 15(1) of this Directive for reproductions and extractions of lawfully accessible works and other subject matter for the purposes of text and data mining.
Article 4 is the general exception and it is open to commercial actors, which makes it the strongest unlicensed-mining argument in the EU. Two limits matter: Article 4(2) permits retention only as long as necessary for mining, and Article 4(3) lets a rights holder opt out.
Directive 96/9/EC Article 7(1)
Member States shall provide for a right for the maker of a database which shows that there has been qualitatively and/or quantitatively a substantial investment in either the obtaining, verification or presentation of the contents to prevent extraction and/or re-utilization of the whole or of a substantial part, evaluated qualitatively and/or quantitatively, of the contents of that database.
This is the provision most often missing from a data deal. A compiled corpus can carry its own right, separate from any copyright in the items inside it. A license that grants "the content" but not the compilation may leave you unable to redistribute the dataset.
Who it applies to
Licensing is the question for anyone who wants to use material they do not own. It is not limited to large models. A 50-hour speech corpus assembled from audiobooks raises the same analysis as a web-scale text crawl, just at a smaller scale and with a smaller damages base.
The question is territorial because the rights are. Copyright is national, and the EU text-and-data-mining exceptions exist only in the member states that implemented them into national law. A US company copying EU-protected content is making reproductions governed by the law of the place where that content is protected, and a US fair-use argument has no EU counterpart, because EU exceptions are a closed list and fair use is not on it. A project touching both sides will usually need a license even where one side alone might have been arguable.
It also applies to material that does not look like a work. A dataset assembled from public-domain recordings can still be somebody's protected compilation, and a marketplace listing rarely resolves which right is being sold. Buyers who read "public domain" as "public domain dataset" are the ones who find this out during diligence rather than before it.
- A dataset you assemble from licensed sources: the license is the whole of your rights, so its scope is your ceiling.
- A dataset you buy from a supplier: you inherit whatever the supplier held, including any gap in their chain.
- A dataset you take from a marketplace: platform terms usually disclaim the rights question, so the analysis falls back to you.
What it costs to get wrong
There is no single fine for training without a license, which is what makes the exposure hard to price. The liability is civil, and it scales with the size of the corpus rather than with the size of the company.
In the US, statutory damages under 17 U.S.C. §504(c) run from 750 to 30,000 dollars per work infringed, and up to 150,000 dollars per work where the infringement is willful. A rights holder does not sue over a corpus as a single object — the works inside it are counted individually, which is how a modest dataset becomes a large number. Attorney's fees are recoverable separately under 17 U.S.C. §505.
In the EU the remedies are national and vary, but the pattern is injunctions plus damages, and courts have now gone both ways on mining. The Hamburg Regional Court held in 2024 that the LAION dataset's reproduction for research purposes fell inside the TDM exception, while a Munich court held in 2026 that unlicensed use of copyrighted music in AI training was infringing. The exception is real. So is the liability when you are standing outside it.
The cost that never appears in a judgment is the one buyers actually carry: a model trained on material you cannot show a license for is a model that will not clear an enterprise procurement review, and the remedy is retraining rather than a settlement.
How to comply when you are buying data
The work is done before signature. A license is worth what it names, and the clauses that decide whether a deal works are the ones most often left vague.
Ask for these five things in writing, and treat a vague answer as a gap you are accepting:
There is a second layer on speech specifically, because a voice recording implicates the speaker as well as any copyright in the recording. Consent documentation is a separate file from the copyright license and is usually the first thing a buyer's counsel asks to see. We build that chain for every hour we source and can produce it on request.
We do not write licenses and nothing here is legal advice. What we can deliver is the source inventory, the consent documentation and a written statement of the rights position, in a form your counsel can review — and we decline projects where that chain cannot be made sound.
- Establish which regime you are actually in. If the material is EU-protected, check whether either TDM exception reaches you. The research exception will not, and the general exception depends on a rights reservation you would have had to verify. If you cannot point to an exception, you need a license.
- Make the grant name the rights. "All rights necessary" is not a grant. Ask specifically for the reproduction right, the right to create derivative works in the form of model weights, and the right to distribute the resulting model to your own customers.
- Ask about the database right separately. If the corpus is a compilation the supplier assembled, the grant should cover the compilation and not only its contents, or you may not be able to pass the dataset on at all.
- Settle sublicensing and model distribution in writing. The commercial question is rarely "may we train". It is "may our customers keep using what we trained", and that is a different grant, usually priced differently.
- Get the source list. A license covering "the dataset" is only as good as the supplier's ability to show what is in it. Ask for the source inventory before you sign, not after a claim arrives.
Related compliance topics
-
AI Training Data and Copyright
ai training data and copyright
-
Voice Data under GDPR
gdpr voice data
-
GDPR Consent for Voice Recording
gdpr voice recording consent
-
Is Voice Personal Data?
is voice personal data
-
Is Voice Biometric Data?
is voice biometric data
-
Biometric Privacy Laws
biometric privacy laws
Not legal advice
We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.
Sourcing data under AI Training Data Licensing?
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.