Glossary
The terms that come up in almost every specification conversation, explained without assuming you already work in the field. Each entry also covers why the term matters when you are the one buying. The list covers speech data and, further down, the vocabulary of embodied AI and robot data collection.
-
ASR
Automatic Speech Recognition — the task of turning recorded speech into text.
-
TTS
Text-to-Speech — the task of generating spoken audio from written text.
-
WER
Word Error Rate — the standard accuracy metric for speech recognition. Lower is better.
-
Transcription
Writing down what is said in a recording. It is the core annotation task in any speech dataset.
-
Annotation
Attaching machine-readable labels to raw data — a transcript, an intent, a speaker identity.
-
Speaker Identification
Identifying who is speaking from the voice itself. Also called voice biometrics.
-
Diarization
Marking which speaker said which segment in a multi-speaker recording.
-
SNR
Signal-to-Noise Ratio — how much louder the speech is than the background noise, measured in decibels (dB).
-
Code-Switching
Switching between two or more languages inside a single sentence. Hinglish and Taglish are the standard examples.
-
Phoneme
The smallest unit of sound in a language that distinguishes one word from another.
-
Prosody
Intonation, rhythm, stress and pauses — how something was said, not just what was said.
-
Ground Truth
Human-verified labels good enough to serve as the training target.
-
Inter-Annotator Agreement
Having different people annotate the same batch and measuring how far apart they land. It is the test of whether a guideline is genuinely unambiguous.
-
PII
Personally Identifiable Information — anything that can identify a specific individual, directly or indirectly.
-
Consent
The speaker knows how their recording will be used, and has signed to say so.
-
Data Compliance
Where the data came from, whether it may be used that way, and whether there is paper to prove it. That whole chain is compliance.
-
GDPR
The EU data protection regulation, which places strict limits on how personal data may be collected and used.
-
Sample Rate
How many times per second audio is captured, in hertz (Hz). Speech data is usually delivered at 16kHz or 48kHz.
-
Script Split
One spoken language written in two different scripts — Hindi in Devanagari and Urdu in Perso-Arabic is the standard case.
-
Low-Resource Language
A language with very little public corpus, so a model cannot learn it from what already exists.
-
Robot Manipulation Data
Robot Manipulation Data — episodes of a robot arm moving objects, recorded together with the commands that produced the motion.
-
Grasping Data
Grasping Data — images, depth maps or point clouds of objects, annotated with the gripper poses that would pick each one up.
-
Dexterous Hand Data
Dexterous Hand Data — demonstrations for multi-fingered robot hands, including in-hand manipulation that a two-finger gripper cannot do.
-
Mobile Robot Navigation Data
Mobile Robot Navigation Data — sensor logs of a robot moving through space, with the poses, maps and trajectories needed to train or evaluate navigation.
-
Humanoid Robot Data
Humanoid Robot Data — whole-body demonstrations and sensor logs from robots with legs and arms, where balance is part of every task.
-
Egocentric Video Data
Egocentric Video Data — first-person footage, usually from a head-mounted camera, showing hands and objects from the wearer's own viewpoint.
-
Tactile Sensor Data
Tactile Sensor Data — contact measurements from a robot's fingertips or skin: pressure, force and, in optical sensors, images of the contact itself.
-
Teleoperation
Teleoperation — a human drives the robot directly, and the robot's motion is recorded as a training demonstration.
-
Motion Capture Data
Motion Capture Data — human body motion recorded as joint trajectories, used for animation and as a reference for robots that imitate people.
-
Sim-to-Real Transfer
Sim-to-Real Transfer — training a policy in simulation and getting it to work on physical hardware, across the gap between the two.
-
Robot Data Augmentation
Robot Data Augmentation — generating additional training episodes from existing ones by transforming images, goals or scene configurations.
-
Robot Data Annotation
Robot Data Annotation — attaching labels to robot episodes: what the task was, whether it succeeded, and when key events happened.
-
Vision-Language-Action Model
Vision-Language-Action Model — a model that takes camera images and a written or spoken instruction and outputs robot actions directly.
-
Warehouse Robotics Data
Warehouse Robotics Data — sensor and telemetry data from robots working in fulfillment centers: picking, moving totes and building pallets.
-
Robot Gripper Data
Robot Gripper Data — episodes and signals tied to the end effector: width, force, vacuum pressure and whether the grasp held.
-
Embodied AI Data
Embodied AI Data — training data for systems that act in the physical world, where every sample records what was perceived, what was done and what happened.
-
Data Licensing
Data Licensing — granting the right to use data you control, under terms that say what the other party may do with it, as distinct from selling it outright.
-
Anonymization
Anonymization — removing or altering identifying information so a person can no longer be identified, in principle irreversibly, which is what separates it from pseudonymization.
-
Biometric Data
Biometric Data — data derived from a person's physical or behavioral characteristics, such as voice or face, which most privacy laws treat as a special and more tightly restricted category.
-
Data Provenance
Data Provenance — the documented history of where data came from, who consented to it and how it was handled, which is what a buyer's legal review asks for before anything can be purchased.
-
DROID Dataset
DROID Dataset — a large in-the-wild robot manipulation dataset collected by a multi-lab consortium, notable for scene diversity rather than task repetition.
-
Open X-Embodiment
Open X-Embodiment — a pooled collection of robot datasets spanning 22 different robot embodiments, built to test whether data from one robot helps another.
-
LeRobot
LeRobot — an open-source library and dataset format for robot learning, which fixes how episodes, video and state are stored so datasets can be shared.
-
LIBERO
LIBERO — a simulation benchmark for robot manipulation built to measure knowledge transfer, with four task suites that isolate different kinds of generalization.
-
AgiBot World
AgiBot World — a large real-world robot dataset collected on a standardized humanoid platform, notable for being recorded at production scale by one organization.
-
Imitation Learning
Imitation Learning — teaching a policy by showing it demonstrations rather than by scoring its attempts, which is why demonstration data is the input it needs.
-
World Model
World Model — a learned model that predicts how a scene will change in response to actions, used both for planning and for generating training data.
Submit a sourcing request
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.