Licensing terms for robot data: the three parties and what each one needs
A robot episode carries rights from the platform, the site, and the operator. The clauses that get skipped are the ones that decide what can be done with a trained model later.
Why a robot episode has more than one owner
A speech recording has one obvious rightsholder and one consent to obtain. A robot episode has a chain. The robot platform contributes software and firmware, and its license may say something about data recorded through its stack. The site contributes the room, the furniture, and the objects, which in a commercial setting can be competitive information. The operator contributes the demonstration itself, and any person visible or audible in the frame contributes a personal data question.
Each link needs its own document before the data is sellable, and the documents have to agree with each other. A facility release that permits internal research does not permit resale, and an operator agreement that assigns the demonstration does not cover the operator voice on the audio track. The order to fix this in is site, then operator, then platform, because the site is usually the slowest to obtain.
The rights that get carved up
The grant clause is where the value is decided, and it is usually written too briefly.
- Training right. Which models, how many, and whether the right survives termination for models already trained.
- Redistribution right. Whether the buyer may pass the data to a contractor, a subsidiary, or a cloud provider for processing. This is refused more often than buyers expect, and the refusal often conflicts with how the buyer actually works.
- Sublicense. Rarely granted, and usually what a marketplace actually needs.
- Derivative treatment. Whether a model trained on the data is treated as a derivative. This is unsettled as a matter of law, so the agreement should state which way the parties treat it rather than leaving it to be argued after training.
- Exclusivity. Whether the same episodes can be licensed to a competitor, and for how long.
- Publication. Whether examples may appear in a paper, a demo, or a customer presentation.
What the platform terms do and do not cover
Platform agreements typically cover the software and the robot itself, not the dataset a collector assembles, and they often restrict redistribution of anything that embeds proprietary logs or firmware output. Two questions are worth asking before a collection starts: does the platform agreement permit selling data recorded on this hardware, and does it require attribution or a separate commercial agreement for that use?
The teleoperation stack carries the same risk in a less obvious place. Motion capture gloves, VR hand-tracking runtimes, and third-party SDKs are frequently licensed for non-commercial use, and a restriction that applies to the tool applies to the dataset produced through it. Read the tool licenses before the collection rather than before the sale.
The site side is about more than permission
A facility release needs a named signer with authority over the space, and a scope that names training use explicitly. The reason is not only legal: video of a working warehouse reveals shelf layout, product mix, and process, and that is information the site owner may consider confidential independently of whether anyone is identifiable.
Objects are a separate matter. A branded package, an unreleased product, or a patented mechanism can all appear in an episode and each brings its own constraint, which is why the practical approach is to inventory the objects in the scene and clear them, rather than to write a broad clause and hope. Where a person can plausibly appear or be heard, the consent question is the same one a field recording has, and it is worth resolving with the same process.
The clauses that get left out
Most disputes about data licensing come from clauses that nobody thought to write, and the list is short enough to check in an afternoon.
- Who owns the annotation layer. Annotations are separately protectable and often produced by a third party under their own terms.
- What happens on termination. Whether models already trained may keep being used and distributed, and whether the buyer must delete the raw episodes.
- Whether the seller may use feedback, corrections, or fine-tuned derivatives for its own model improvement.
- Whether the buyer may publish evaluation results, since a benchmark number can disclose more than the data itself.
- The acceptance process and the remedy when a delivered batch fails it.
- Governing law and venue, which decides how expensive the rest of the document is to enforce.
The schedule that prevents most of it
Attach a one-page schedule to the agreement and make it the operative description: data description and version, delivery format, the rights granted and the rights withheld, the parties and their roles, the site and consent documents by reference, annotation ownership, and the training-versus-redistribution split. Anything not on the schedule is a future argument, and the schedule is the document that a reviewer, an auditor, or an acquiring company will actually read.
Keep it current. A dataset that grows a second collection batch, a second site, or a second annotation vendor has changed in ways the schedule has to reflect, and a schedule that describes a dataset that no longer exists is worse than no schedule at all.