Running a teleoperation shift: rig setup, operator training, and the data you throw away

Leader-follower rigs, operator training, session length, and the failure patterns that show up in raw teleoperation data before any filtering.

Teleoperation is still the default way to get manipulation demonstrations. A person drives the robot, the robot records its own state, and the result is a trajectory a policy can imitate. It is simple to describe and hard to run well, because the quality of the output is decided by the rig and by the operator's habits long before anyone opens a training script.

Match the leader to the follower's kinematics

Joint-mapped rigs copy the leader's joint angles straight to the follower with no inverse kinematics in the loop. The ALOHA work uses this arrangement, with a low-cost leader arm driving a matching follower, and the GELLO framework does the same with 3D-printed arms built to share the follower's kinematics. Cartesian rigs are the other family: the operator drives the end effector with a spacemouse or a handheld tracker, and the arm solves inverse kinematics every control step. That is more flexible across different arms, and it produces smoother end-effector paths with joint commands that can look unnatural near singularities.

Decide explicitly what the recorded action is. Joint positions, joint velocities, and end-effector poses are all defensible, but a dataset that mixes them is not. Record the commanded action and the achieved state separately, because their difference is the only clean way to diagnose tracking error later. Measure latency instead of assuming it: a joint-mapped pair can sit under a few tens of milliseconds, while a wireless link plus an inverse-kinematics solve can be several times that.

Train the operator like a pilot

An operator who has never done this will produce a first block of data that is not worth keeping: exploratory motion, uneven timing, grasp points that wander. Budget a practice period and treat it as part of the project cost. A workable protocol is to repeat a mock task until episode duration stops falling and the start pose stops drifting, usually somewhere between thirty and sixty episodes, and only then start recording the real set.

Write the verbal protocol down before the first session: when to abort an episode rather than push through, where the arm returns between episodes, how the object is held so the camera keeps it in view, and whether the operator narrates. Those are otherwise decided differently on different days. Fatigue shows up as slower motion and more re-grasps rather than obvious mistakes, so schedule the shift in blocks with real breaks.

Session length is set by reset time

Reset cost dominates a shift. If a demonstration takes eight seconds and resetting the scene takes forty, most of the shift is scene setup. Fix that before adding operators: pre-positioned object trays, spare objects staged just outside the camera frame, and a single defined reset pose the arm returns to every time.

Record episodes with clean boundaries. Start only after the arm is at the reset pose and the scene is final, and stop before the operator reaches in to reset. Setup and reset motion inside an episode is one of the most common defects in raw teleoperation data. Plan throughput per operator per hour after warm-up rather than a per-day target: a good operator produces a few hundred episodes in a shift across a task family, not thousands.

What bad teleoperation data looks like

These patterns account for most rejected episodes, and all of them are visible before annotation starts.

  • Hesitation pauses mid-trajectory. The operator stops to decide, the recorded action holds still, then resumes with a jerk.
  • Aborted and restarted attempts inside one episode. A missed grasp, a lift, a second try, and all of it in one file with one label.
  • Leader-follower tracking loss. The follower lags or hits a joint limit while the leader keeps moving, so the recorded state no longer matches the commanded action.
  • Rig occlusion. A wrist camera blocked by the arm or the operator's hand, so the observation stream goes dark where the interesting motion is.
  • Exposure shifts when the operator's body blocks a light source. The frames are valid, but the visual distribution no longer matches deployment.

Filter mechanically before anyone watches

Run an automated pass at ingest, before annotation, and reject on numbers rather than impressions: stationary runs longer than a threshold, joint velocity spikes above the hardware limit, episode length outliers against the task median, and command-to-state divergence above a tolerance. Then review the flagged episodes at speed. Keep every reject with a reason code, because a reject log broken down by operator and by day tells you whether the problem is the person, the rig, or the task, which is the difference between a fix and another week of the same defect.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com