Data categories

Structure here is open — the catalog can grow to whatever a requirement needs. What is fixed is the template: every category page explains what the data is, what buyers get wrong about it, and which specification detail most often decides whether a delivery is usable.

Speech data

The categories we produce today. Everything here can be specified, sourced and delivered. 18 categories

  • Call Center Speech

    Recorded customer-service calls, real or simulated, capturing both the agent and the caller side, used for call center QC, intent recognition, and dialogue systems.

  • Conversational Speech

    Speech from natural conversation between two or more people, on open or semi-structured topics, used for conversational AI, voice assistants, and small talk models.

  • Singing Voice

    Vocal recordings with melody, in both a cappella and accompanied form, used for singing voice synthesis, music information retrieval, and lyric alignment.

  • Speech Commands

    Targeted recordings of short command words or phrases, usually with many speakers reading each entry several times over, used for wake words and on-device command recognition.

  • Read Speech

    Recordings of speakers reading specified text, with clear pronunciation and known text, the base material for TTS and ASR.

  • Podcast Speech

    Long-form podcast and interview audio, either solo monologue or two-person conversation, used for long-form speech recognition and speaker modeling.

  • Multilingual Speech

    Speech data covering multiple languages within one project, used for multilingual ASR, cross-lingual transfer, and language identification.

  • Noisy Speech

    Speech collected under background noise — street, in-car, restaurant, office and other real environments — used for noise-robust models and speech enhancement.

  • Code-Switching Speech

    Speech that mixes two or more languages inside a single utterance — Hinglish, Spanglish, Taglish — used for recognition in real spoken settings.

  • Children Speech

    Speech data from child speakers, grouped by age band, used for children's speech recognition and children's education products.

  • Audiobook

    Continuous long-form narration in audiobook style, steady in intonation and generous in duration, the main source material for high-quality TTS.

  • Short Video Speech

    Speech in the talking-head style of short video — fast, emotionally strong, colloquial — used for short video subtitles and content understanding.

  • Live Stream Speech

    Long-session spoken content from live streams — host monologue and responses to audience interaction — fast-paced and heavily improvised.

  • Voice Assistant

    Command and dialogue data for voice assistants, usually with intent annotation, covering the full chain from wake word to understanding to response.

  • Elderly Voice

    Speech data from elderly speakers, usually stratified by age band and health status, used for voice products built for older users.

  • Accented English

    English speech from non-native speakers or specific regional accents, grouped by accent origin, used to improve accent robustness in ASR.

  • Speech Translation

    Parallel data pairing source-language audio with target-language translation, used for speech-to-text and speech-to-speech translation models.

  • In-the-Wild Speech

    Speech collected under fully natural conditions — no studio, no topic constraints — as close to real usage as collection gets.

New data paradigms

Embodied AI, synthetic data and world models — categories where the data formats are still being standardized. 3 categories

  • Embodied AI Data

    Multimodal data generated as robots and embodied agents operate in real environments — vision, action trajectories, and language instructions.

  • Synthetic Data

    Data generated by models or simulation engines, used to fill gaps where real data is scarce, cover long-tail scenarios, and control cost.

  • World Model Data

    Environment interaction data for training world models, emphasizing temporal consistency, physical plausibility, and predictability.

Non-speech data

Adjacent categories we source through partner producers. Same sourcing model, different production chains. 15 categories

  • LLM Training Data

    Text corpora for LLM pretraining and fine-tuning — instruction data, preference data, and multi-turn dialogue data.

  • Computer Vision Data

    Annotated image and video data, with annotation types including classification, bounding boxes, segmentation masks, and keypoints.

  • NLP Data

    Annotated text for NLP tasks — named entities, dependency syntax, sentiment polarity, and relation extraction.

  • Multilingual Language Data

    Parallel or non-parallel text corpora spanning multiple languages, used for multilingual models and low-resource language research.

  • Translation Data

    Parallel corpora pairing source and target language, used for machine translation training and evaluation.

  • Dictionary Data

    Structured dictionary entries with definitions, parts of speech, pronunciation, and example sentences, used for dictionary products and lexical understanding tasks.

  • Text Data

    General-purpose text corpora from web pages, forums, reviews and similar sources, used for language model pretraining and text classification.

  • Video Data

    Annotated video footage, with the annotation needed for tasks such as action recognition, object tracking, and temporal segmentation.

  • Satellite Imagery Data

    Remote-sensing satellite or aerial imagery, with multispectral bands and geographic annotation, used for land-cover classification, change detection, and agricultural monitoring.

  • First-Person View Data

    Video and sensor data captured from the operator's viewpoint with wearable devices, used for embodied AI and action understanding.

  • Autonomous Driving Data

    Road-scene data captured by vehicle sensors — cameras, lidar, and annotation boxes — used to train perception models.

  • Legal Data

    Legal filings, case law, and statutory text, used for legal search, contract analysis, and legal question-answering models.

  • Patent Data

    Patent applications, claims, and citation data, used for patent search, technology trend analysis, and infringement risk detection.

  • News Data

    News articles and reportage with date, source, and topic tags, used for event extraction and temporal analysis.

  • Academic Data

    Academic papers, abstracts, and citation data, used for research question answering, literature review, and knowledge graph construction.

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com