Embodied AI Data Collection

Collection for robots and physical agents, scoped as a task matrix rather than an hour count. What the project involves, what moves the cost, and what lands in your hands at delivery.

What the project actually involves

A collection engagement runs in four stages, and the first one is the longest. Specification means writing down the task list as verb plus object, the robot platform, the sensor configuration, the environments, and what counts as a completed attempt. This is done before any hardware is booked, because a change to the task list after production starts is a change to everything downstream.

The pilot is a short run — a day or two of collection in one environment, using the exact capture setup and annotation conventions intended for production. The buyer reviews the episodes and either signs off or changes the specification while the cost of changing it is still small.

Production is the scheduled work: sessions in each environment until the episode targets are met, with the same operators, the same calibration procedure, and the same reset discipline throughout. QA runs alongside production rather than at the end, because a stream-alignment fault found on day two is a fix, and the same fault found at delivery is a re-run.

What decides the cost

Embodied collection is priced by time-on-task, not by file size, and time-on-task is dominated by everything around the recording. A teleoperated arm needs an operator at the rig, a scene that is reset between attempts, objects put back in position, and batteries swapped. The fraction of a session day that becomes usable episode time is often well under half.

  • Sensor count — each added camera or depth sensor multiplies data volume, adds synchronization work, and requires its own calibration. The third camera usually costs more in engineering than in hardware.
  • Teleoperation versus autonomous capture — teleoperation is bound by operator hours; autonomous capture can run longer but yields fewer successful episodes per hour and needs a safety plan for every environment.
  • Annotation depth — raw trajectories, episode segmentation, action labels, and language instructions are four different price points on the same recording. Instruction annotation is usually the largest single line.
  • Scene diversity — ten environments cost more than ten repetitions in one environment, and they are the part that decides whether the model generalizes beyond the room it was recorded in.

How long an engagement takes

For a common research platform and a task list that is already written, a pilot can usually start within a few weeks. Production then depends almost entirely on the episode target and the environment count: a few hundred manipulation episodes across three environments is a matter of weeks, while a project spanning dozens of scenes and object categories runs for months.

Two factors stretch timelines more than buyers expect. The first is hardware availability — if the project needs a platform nobody has standing access to, the hardware has to be located or shipped, and that is calendar time nobody controls. The second is operator skill: teleoperation quality varies enormously between operators, and recruiting and screening them is part of the schedule, not an afterthought.

QA and packaging add one to two weeks on top of production, and they are not compressible without cutting the checks that catch delivery failures.

What you receive, and what you decide first

A delivery contains the episodes themselves with aligned sensor streams, a calibration record for every session, the annotation schema the data was labeled under, and a QA report describing what was checked and what was excluded. Failure episodes are included or excluded according to the success criteria agreed at specification, and either choice is documented.

Before work starts, the buyer has to settle the platform and its action-space convention, freeze the task list, decide the environment spread and object inventory, set the annotation depth, and state whether failed attempts are retained. These are the decisions that cannot be made late — each one changes the pilot, and the pilot is what the rest of the project is built on.

Questions we get asked

Can you collect on our own robot platform?

Sometimes. If the platform is a widely used research model, facilities that already run it are the fastest route. If it is your own prototype, the realistic options are shipping the hardware to a collection site, or having your own team operate it while we handle specification, annotation, and QA.

Do we get raw sensor streams or processed episodes?

Both are possible, and the choice changes the price. Raw streams with calibration are cheaper; segmented episodes with action labels cost more because segmentation and labeling are human work. Most buyers take processed episodes and keep the raw streams as delivered files.

How is a pilot batch defined?

A pilot is a small slice of the real specification — the same platform, the same sensors, the same environments where possible, and the full annotation depth on a limited episode count. Its purpose is to expose disagreements between the specification and reality while they are still cheap to fix.

Keep reading

  • AI Data Brokerage

    A data broker sits between the people who need training data and the people who can produce it. We source to order — no inventory, no license resale, no recycled datasets.

  • AI Training Data Providers

    What to check before you sign with a training data provider — and how we compare on each point.

  • AI Training Data Marketplace

    Marketplaces sell what already exists. We source what you actually need. Here is when each model makes sense.

  • Full data catalog

    36 data categories across 120 languages, and how to specify each one.

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com