Model rights and data rights are two different assets
Buying data does not buy the model, and licensing a model does not carry the data rights with it. The boundary questions that have to be written into the contract.
Why the two get confused
A dataset licence and a model licence are different instruments, and each one says almost nothing about the other. Teams get caught because the two are discussed in the same conversation and the boundary between them is never written down.
The mistake runs in both directions. A team licenses a corpus and assumes the resulting model is theirs to distribute, then finds the grant covers training and internal evaluation. A team licenses a model and assumes a warranty about the training data comes with it, then finds the model licence carries no representation about the data at all — which leaves the data risk sitting with the licensee, who is the one shipping a product.
The four boundary questions
Does the data grant cover the model's outputs? Generated audio or text is not a copy of the training material, and it is not obviously a derivative of it either. Because the default answer is unsettled, it has to be allocated by the contract; a silent clause is a clause that will be argued about later, usually after the product has shipped.
Can the model be distributed? Training, internal evaluation and shipping a product are three separate permissions. The third has to be explicit, and it has to match how you actually distribute: an API, an on-device build and a hosted service each raise a different question about whether the end user receives something derived from the licensed material.
Can the vendor use your model or its outputs? If the vendor fine-tunes on your corpus, or uses your product's outputs to improve their own tooling, they are building an asset out of your material. Ask whether the agreement permits it, and if it does, whether it is confined to their internal quality work.
What do you get from the vendor's tooling? If annotations were produced by a system the vendor built, you may be receiving labels from a tool you have no rights to and no visibility into. That is usually acceptable — but the contract should say so rather than leave it to be discovered during diligence.
The synthetic data loop
Synthetic material generated from a licensed corpus is a boundary case that catches teams out. Three questions decide it. Does the grant permit training on synthetic material derived from the licensed data? Who owns the synthetic set — you, the vendor, or nobody in particular? And may the vendor sell that synthetic set to anyone else?
The third is the one to press on. A vendor who generates synthetic data from your licensed corpus and sells the result has effectively sold the same information twice, in a form your exclusivity may not reach. If you paid for exclusivity or for a fresh collection, the derivation rule belongs in the contract: synthetic material derived from the corpus sits inside the exclusivity scope, and the vendor may not produce it for other clients.
Rights to improvements
When a vendor works on your data, they learn things — which screening criteria work, which annotation conventions hold up, how a particular language behaves. That knowledge is not an asset anyone can own, and trying to own it is a losing fight that wastes the negotiation.
What can be owned is the artefacts. Draw the line in writing: improvements to the vendor's tools and methods stay with the vendor; the trained model, the annotations and the reports produced for your project stay with you. Then add the clause that actually protects you — the vendor may not use your corpus, your roster or your outputs to build a product for another client in your market.
A one-page allocation to attach
List every asset on a single page with three columns: who owns it, who may use it, and what the counterparty must not do with it. The assets are raw recordings, performer grants, transcripts and labels, roster and provenance records, the specification and guideline documents, the vendor's tooling, any model trained under the agreement, that model's outputs, and any synthetic material derived from the corpus.
The discipline of writing it out is the point. Most disputes in this area are not about who is right on the law; they are about which asset a clause was written for. A page that names every asset, signed alongside the contract, removes most of that ambiguity at the cost of one meeting.
None of this substitutes for legal advice. The allocation questions above are settled differently in different jurisdictions, and a lawyer should look at the actual grant language before it is signed.