What is Vision-Language-Action Model?

Vision-Language-Action Model — a model that takes camera images and a written or spoken instruction and outputs robot actions directly.

The core idea is that actions become another language. RT-2, built on a 55-billion-parameter vision-language backbone, tokenizes each action as text and trains on robot episodes and web data together — which is where its ability to follow instructions that never appeared in robot data comes from. OpenVLA is the open version: 7 billion parameters, trained on about 970,000 episodes from Open X-Embodiment, fine-tunable with LoRA by updating roughly 1.4% of its parameters.

For data buyers the unit of value changes. A VLA does not need more hours of one task; it needs episodes that pair an instruction with the right behavior across many objects, scenes and robot bodies — the reason Open X-Embodiment pools 22 embodiments rather than scaling one. Failure episodes matter more here than anywhere else, because the model has to learn what an instruction does not mean.

There is also a rate mismatch to plan around: large VLAs run closed-loop control at only a few hertz, so demonstrations recorded at 50 Hz are downsampled or chunked before training, and the chunking convention has to match the deployment. Teams training a VLA usually specify data at the episode level — embodiment mix, instruction style, failure coverage — and collection or conversion is built to that specification rather than sold by the hour.

Related terms

  • Robot Manipulation Data

    Robot Manipulation Data — episodes of a robot arm moving objects, recorded together with the commands that produced the motion.

  • Egocentric Video Data

    Egocentric Video Data — first-person footage, usually from a head-mounted camera, showing hands and objects from the wearer's own viewpoint.

  • Sim-to-Real Transfer

    Sim-to-Real Transfer — training a policy in simulation and getting it to work on physical hardware, across the gap between the two.

Keep reading

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com