Glossary

The terms that come up in almost every specification conversation, explained without assuming you already work in the field. Each entry also covers why the term matters when you are the one buying. The list covers speech data and, further down, the vocabulary of embodied AI and robot data collection.

  • ASR

    Automatic Speech Recognition — the task of turning recorded speech into text.

  • TTS

    Text-to-Speech — the task of generating spoken audio from written text.

  • WER

    Word Error Rate — the standard accuracy metric for speech recognition. Lower is better.

  • Transcription

    Writing down what is said in a recording. It is the core annotation task in any speech dataset.

  • Annotation

    Attaching machine-readable labels to raw data — a transcript, an intent, a speaker identity.

  • Speaker Identification

    Identifying who is speaking from the voice itself. Also called voice biometrics.

  • Diarization

    Marking which speaker said which segment in a multi-speaker recording.

  • SNR

    Signal-to-Noise Ratio — how much louder the speech is than the background noise, measured in decibels (dB).

  • Code-Switching

    Switching between two or more languages inside a single sentence. Hinglish and Taglish are the standard examples.

  • Phoneme

    The smallest unit of sound in a language that distinguishes one word from another.

  • Prosody

    Intonation, rhythm, stress and pauses — how something was said, not just what was said.

  • Ground Truth

    Human-verified labels good enough to serve as the training target.

  • Inter-Annotator Agreement

    Having different people annotate the same batch and measuring how far apart they land. It is the test of whether a guideline is genuinely unambiguous.

  • PII

    Personally Identifiable Information — anything that can identify a specific individual, directly or indirectly.

  • Consent

    The speaker knows how their recording will be used, and has signed to say so.

  • Data Compliance

    Where the data came from, whether it may be used that way, and whether there is paper to prove it. That whole chain is compliance.

  • GDPR

    The EU data protection regulation, which places strict limits on how personal data may be collected and used.

  • Sample Rate

    How many times per second audio is captured, in hertz (Hz). Speech data is usually delivered at 16kHz or 48kHz.

  • Script Split

    One spoken language written in two different scripts — Hindi in Devanagari and Urdu in Perso-Arabic is the standard case.

  • Low-Resource Language

    A language with very little public corpus, so a model cannot learn it from what already exists.

  • Robot Manipulation Data

    Robot Manipulation Data — episodes of a robot arm moving objects, recorded together with the commands that produced the motion.

  • Grasping Data

    Grasping Data — images, depth maps or point clouds of objects, annotated with the gripper poses that would pick each one up.

  • Dexterous Hand Data

    Dexterous Hand Data — demonstrations for multi-fingered robot hands, including in-hand manipulation that a two-finger gripper cannot do.

  • Mobile Robot Navigation Data

    Mobile Robot Navigation Data — sensor logs of a robot moving through space, with the poses, maps and trajectories needed to train or evaluate navigation.

  • Humanoid Robot Data

    Humanoid Robot Data — whole-body demonstrations and sensor logs from robots with legs and arms, where balance is part of every task.

  • Egocentric Video Data

    Egocentric Video Data — first-person footage, usually from a head-mounted camera, showing hands and objects from the wearer's own viewpoint.

  • Tactile Sensor Data

    Tactile Sensor Data — contact measurements from a robot's fingertips or skin: pressure, force and, in optical sensors, images of the contact itself.

  • Teleoperation

    Teleoperation — a human drives the robot directly, and the robot's motion is recorded as a training demonstration.

  • Motion Capture Data

    Motion Capture Data — human body motion recorded as joint trajectories, used for animation and as a reference for robots that imitate people.

  • Sim-to-Real Transfer

    Sim-to-Real Transfer — training a policy in simulation and getting it to work on physical hardware, across the gap between the two.

  • Robot Data Augmentation

    Robot Data Augmentation — generating additional training episodes from existing ones by transforming images, goals or scene configurations.

  • Robot Data Annotation

    Robot Data Annotation — attaching labels to robot episodes: what the task was, whether it succeeded, and when key events happened.

  • Vision-Language-Action Model

    Vision-Language-Action Model — a model that takes camera images and a written or spoken instruction and outputs robot actions directly.

  • Warehouse Robotics Data

    Warehouse Robotics Data — sensor and telemetry data from robots working in fulfillment centers: picking, moving totes and building pallets.

  • Robot Gripper Data

    Robot Gripper Data — episodes and signals tied to the end effector: width, force, vacuum pressure and whether the grasp held.

  • Embodied AI Data

    Embodied AI Data — training data for systems that act in the physical world, where every sample records what was perceived, what was done and what happened.

  • Data Licensing

    Data Licensing — granting the right to use data you control, under terms that say what the other party may do with it, as distinct from selling it outright.

  • Anonymization

    Anonymization — removing or altering identifying information so a person can no longer be identified, in principle irreversibly, which is what separates it from pseudonymization.

  • Biometric Data

    Biometric Data — data derived from a person's physical or behavioral characteristics, such as voice or face, which most privacy laws treat as a special and more tightly restricted category.

  • Data Provenance

    Data Provenance — the documented history of where data came from, who consented to it and how it was handled, which is what a buyer's legal review asks for before anything can be purchased.

  • DROID Dataset

    DROID Dataset — a large in-the-wild robot manipulation dataset collected by a multi-lab consortium, notable for scene diversity rather than task repetition.

  • Open X-Embodiment

    Open X-Embodiment — a pooled collection of robot datasets spanning 22 different robot embodiments, built to test whether data from one robot helps another.

  • LeRobot

    LeRobot — an open-source library and dataset format for robot learning, which fixes how episodes, video and state are stored so datasets can be shared.

  • LIBERO

    LIBERO — a simulation benchmark for robot manipulation built to measure knowledge transfer, with four task suites that isolate different kinds of generalization.

  • AgiBot World

    AgiBot World — a large real-world robot dataset collected on a standardized humanoid platform, notable for being recorded at production scale by one organization.

  • Imitation Learning

    Imitation Learning — teaching a policy by showing it demonstrations rather than by scoring its attempts, which is why demonstration data is the input it needs.

  • World Model

    World Model — a learned model that predicts how a scene will change in response to actions, used both for planning and for generating training data.

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com