Copyright diligence before you license a dataset
What to check, what documents to demand, and the red flags that mean walking away — before you sign a license for training data.
Copyright diligence on a dataset is not about reading the license you are offered. It is about the rights underneath that license. A seller can only grant what they hold, and the license you sign cannot fill a gap that exists above it in the chain.
That is the whole shape of the exercise: inventory the material, trace the chain of title, compare the grant to your actual use, and record what you found. The steps below are the working version.
Step one: inventory what is actually in the corpus
Before any rights question, establish what the material is. How many recording sessions, and what kinds of content: spontaneous conversation, read speech, scripted dialogue, interviews. Whether any non-speech material is embedded — music, broadcast audio, ambient recordings. And whether any of the content is text that was read aloud, which is the item most often missed: the recording may belong to the studio while the text belongs to someone else entirely.
The output of this step is a content inventory that maps each category to its own rights question. "Four hundred hours of Spanish" is not an inventory; it is a number.
Step two: trace the chain of title
For each category, identify who held the rights at each transfer, and what the transfer document actually said. The documents to demand:
- The speaker agreement used at collection, in the version that was actually signed — not the current template.
- The contract between the collector and the seller, when those are different parties.
- The upstream license, when the seller acquired the corpus rather than producing it.
- A written statement of which files, if any, are excluded from the grant.
Step three: compare the grant to your actual use
A valid license for the wrong use is still a problem. The questions to answer against the actual product:
- Does the grant name machine learning training, or only "use" and "reproduction"?
- Does it cover model outputs, or only the data itself?
- Does it permit sublicensing to your customers, if your product exposes the model to them?
- Is it perpetual, or does it expire while the model is still in service?
- Does it cover every territory where your users are?
Red flags
Each of these has ended a deal after signing, at a much higher cost than walking away beforehand:
- The seller cannot name the recording period, the collection method, or who did the recording.
- The speaker agreement grants all rights, but was signed by an entity that no longer exists, with no assignment trail.
- The corpus includes audio scraped from user uploads, with no consent records.
- Metadata contradicts the manifest — dates, languages or counts that do not reconcile.
- The seller offers to sort out the paperwork after payment, or says rights for part of the corpus are "in progress."
Record the outcome, not just the decision
Diligence produces a memo: what was reviewed, what was missing, what was accepted with conditions, and what residual risk remains. It feels bureaucratic until the first time a corpus is audited, resold or challenged, and the memo is the only thing that shows the decision was made with open eyes.
One discipline keeps the memo honest: write it before signing, not after. A memo that cannot be written clearly is a signal about the deal, not about the memo.