AI Data Licensing

What the license actually permits is more important than the dataset specification. Here is what to check.

The license is the product

A dataset you cannot legally use is not a dataset. Buyers spend most of their evaluation effort on specification and almost none on licensing, then discover at legal review that the terms prohibit the thing they bought it for.

The common failure is a license that permits research use but not commercial use, or permits internal use but not the distribution of a model trained on the data. Both are easy to miss in a long agreement.

What to check in a data license

These are the clauses that most often decide whether a deal works.

  • Permitted use — research, internal commercial, or product-facing. Model training is not automatically covered by all three.
  • Model distribution rights — whether you may ship a model trained on the data, and whether you owe anything if you do.
  • Exclusivity — whether the same data is available to competitors.
  • Derivative works — whether you may create and distribute modified versions.
  • Term and termination — perpetual, or time-limited with renewal.
  • Sublicensing — whether you may pass the data or the model to a customer or contractor.

The chain behind the license

The license you sign is only as strong as the rights the provider actually holds. If the provider does not have documented consent from the speakers, they cannot grant you rights they do not have.

This is where sourcing to order has an advantage over licensing existing datasets. When data is produced against your order, the consent chain is built for your use case from the beginning rather than retrofitted.

For anything involving children, the chain is stricter — guardian consent is required, and the permitted use should be explicit about it.

Licensing from us

Deliveries come with the consent documentation and collection methodology that support the rights being granted. We do not retain the right to resell your dataset.

If your legal team has specific requirements — jurisdiction, model distribution, sublicensing — raise them before production starts. Terms are easier to build in at the beginning than to add later.

Questions we get asked

Do I own the data after I buy it?

Terms vary by project and are set in the agreement. What is consistent is that we do not keep a right to resell what we produced for you.

Can you provide GDPR-compliant consent documentation?

For projects involving EU data subjects, yes — the consent chain is built to satisfy GDPR requirements, and the documentation is delivered with the dataset.

What about data involving children?

Guardian consent is required and the permitted use is written explicitly. This is a hard requirement; we do not produce children's voice data without it.

Keep reading

  • AI Data Brokerage

    A data broker sits between the people who need training data and the people who can produce it. We source to order — no inventory, no license resale, no recycled datasets.

  • AI Training Data Providers

    What to check before you sign with a training data provider — and how we compare on each point.

  • AI Training Data Marketplace

    Marketplaces sell what already exists. We source what you actually need. Here is when each model makes sense.

  • Full data catalog

    36 data categories across 120 languages, and how to specify each one.

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com