What is Open X-Embodiment?

Open X-Embodiment — a pooled collection of robot datasets spanning 22 different robot embodiments, built to test whether data from one robot helps another.

Open X-Embodiment, released in late 2023, is a pooling effort rather than a new collection: sixty existing robot datasets from many labs were converted into one shared format (RLDS) and released together. The result spans 22 distinct robot embodiments, more than a million real episodes and several hundred distinct skills. The scale is not the interesting part — the embodiment count is.

The question it was built to answer is whether a robot learns better from data collected on other robots. The answer was yes: models trained on the pooled set outperformed models trained only on data from the target robot. That result is why cross-embodiment data is now a standard line item in robot learning budgets, and why "which embodiments are in it" is a more useful question to ask than "how many hours".

Two caveats matter if you are buying or assembling something similar. Pooling heterogeneous sources means pooling heterogeneous conventions — camera placement, action spaces, control rates and success criteria all differ, and the conversion into a shared format is where information quietly goes missing. And the pool leans toward tabletop manipulation with parallel-jaw grippers; if your platform is a humanoid or a dexterous hand, the cross-embodiment benefit is real but thinner.

Related terms

  • DROID Dataset

    DROID Dataset — a large in-the-wild robot manipulation dataset collected by a multi-lab consortium, notable for scene diversity rather than task repetition.

  • Robot Manipulation Data

    Robot Manipulation Data — episodes of a robot arm moving objects, recorded together with the commands that produced the motion.

  • LeRobot

    LeRobot — an open-source library and dataset format for robot learning, which fixes how episodes, video and state are stored so datasets can be shared.

Keep reading

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com