Copyright diligence before you license a dataset

What to check, what documents to demand, and the red flags that mean walking away — before you sign a license for training data.

Copyright diligence on a dataset is not about reading the license you are offered. It is about the rights underneath that license. A seller can only grant what they hold, and the license you sign cannot fill a gap that exists above it in the chain.

That is the whole shape of the exercise: inventory the material, trace the chain of title, compare the grant to your actual use, and record what you found. The steps below are the working version.

Step one: inventory what is actually in the corpus

Before any rights question, establish what the material is. How many recording sessions, and what kinds of content: spontaneous conversation, read speech, scripted dialogue, interviews. Whether any non-speech material is embedded — music, broadcast audio, ambient recordings. And whether any of the content is text that was read aloud, which is the item most often missed: the recording may belong to the studio while the text belongs to someone else entirely.

The output of this step is a content inventory that maps each category to its own rights question. "Four hundred hours of Spanish" is not an inventory; it is a number.

Step two: trace the chain of title

For each category, identify who held the rights at each transfer, and what the transfer document actually said. The documents to demand:

  • The speaker agreement used at collection, in the version that was actually signed — not the current template.
  • The contract between the collector and the seller, when those are different parties.
  • The upstream license, when the seller acquired the corpus rather than producing it.
  • A written statement of which files, if any, are excluded from the grant.

Step three: compare the grant to your actual use

A valid license for the wrong use is still a problem. The questions to answer against the actual product:

  • Does the grant name machine learning training, or only "use" and "reproduction"?
  • Does it cover model outputs, or only the data itself?
  • Does it permit sublicensing to your customers, if your product exposes the model to them?
  • Is it perpetual, or does it expire while the model is still in service?
  • Does it cover every territory where your users are?

Red flags

Each of these has ended a deal after signing, at a much higher cost than walking away beforehand:

  • The seller cannot name the recording period, the collection method, or who did the recording.
  • The speaker agreement grants all rights, but was signed by an entity that no longer exists, with no assignment trail.
  • The corpus includes audio scraped from user uploads, with no consent records.
  • Metadata contradicts the manifest — dates, languages or counts that do not reconcile.
  • The seller offers to sort out the paperwork after payment, or says rights for part of the corpus are "in progress."

Record the outcome, not just the decision

Diligence produces a memo: what was reviewed, what was missing, what was accepted with conditions, and what residual risk remains. It feels bureaucratic until the first time a corpus is audited, resold or challenged, and the memo is the only thing that shows the decision was made with open eyes.

One discipline keeps the memo honest: write it before signing, not after. A memo that cannot be written clearly is a signal about the deal, not about the memo.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com