What is LIBERO?

LIBERO — a simulation benchmark for robot manipulation built to measure knowledge transfer, with four task suites that isolate different kinds of generalization.

LIBERO is a benchmark rather than a data product. It runs in simulation on robosuite and MuJoCo with a Franka arm at a tabletop, and it is organized into four suites that each vary one thing: LIBERO-Spatial moves object positions, LIBERO-Object swaps the objects themselves, LIBERO-Goal changes what the instruction asks for, and LIBERO-Long chains several steps into a single instruction. Around 130 tasks in total, with human demonstrations collected by teleoperation alongside scripted ones.

The design is the contribution. By changing one axis at a time, the suites separate "the model cannot see the object" from "the model cannot follow the instruction" from "the model cannot hold a plan together" — distinctions that a single aggregate success rate hides completely. The Long suite is where most policies fall over, because a multi-step task fails if any step does and the failure rate compounds across them.

For a buyer, LIBERO is most useful as a shared reference point. Because the suites are fixed and public, a success rate on them is comparable between two teams in a way an in-house benchmark never is — which makes it a defensible acceptance criterion for a policy deliverable, provided the brief names which suite and which split. It is not a substitute for real-robot evaluation: a policy that scores well on LIBERO-Long can still fail on the first contact-rich step in the physical world.

Related terms

  • Sim-to-Real Transfer

    Sim-to-Real Transfer — training a policy in simulation and getting it to work on physical hardware, across the gap between the two.

  • Robot Manipulation Data

    Robot Manipulation Data — episodes of a robot arm moving objects, recorded together with the commands that produced the motion.

  • Imitation Learning

    Imitation Learning — teaching a policy by showing it demonstrations rather than by scoring its attempts, which is why demonstration data is the input it needs.

Keep reading

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com