Data categories
Structure here is open — the catalog can grow to whatever a requirement needs. What is fixed is the template: every category page explains what the data is, what buyers get wrong about it, and which specification detail most often decides whether a delivery is usable.
Speech data
The categories we produce today. Everything here can be specified, sourced and delivered. 18 categories
-
Call Center Speech
Recorded customer-service calls, real or simulated, capturing both the agent and the caller side, used for call center QC, intent recognition, and dialogue systems.
-
Conversational Speech
Speech from natural conversation between two or more people, on open or semi-structured topics, used for conversational AI, voice assistants, and small talk models.
-
Singing Voice
Vocal recordings with melody, in both a cappella and accompanied form, used for singing voice synthesis, music information retrieval, and lyric alignment.
-
Speech Commands
Targeted recordings of short command words or phrases, usually with many speakers reading each entry several times over, used for wake words and on-device command recognition.
-
Read Speech
Recordings of speakers reading specified text, with clear pronunciation and known text, the base material for TTS and ASR.
-
Podcast Speech
Long-form podcast and interview audio, either solo monologue or two-person conversation, used for long-form speech recognition and speaker modeling.
-
Multilingual Speech
Speech data covering multiple languages within one project, used for multilingual ASR, cross-lingual transfer, and language identification.
-
Noisy Speech
Speech collected under background noise — street, in-car, restaurant, office and other real environments — used for noise-robust models and speech enhancement.
-
Code-Switching Speech
Speech that mixes two or more languages inside a single utterance — Hinglish, Spanglish, Taglish — used for recognition in real spoken settings.
-
Children Speech
Speech data from child speakers, grouped by age band, used for children's speech recognition and children's education products.
-
Audiobook
Continuous long-form narration in audiobook style, steady in intonation and generous in duration, the main source material for high-quality TTS.
-
Short Video Speech
Speech in the talking-head style of short video — fast, emotionally strong, colloquial — used for short video subtitles and content understanding.
-
Live Stream Speech
Long-session spoken content from live streams — host monologue and responses to audience interaction — fast-paced and heavily improvised.
-
Voice Assistant
Command and dialogue data for voice assistants, usually with intent annotation, covering the full chain from wake word to understanding to response.
-
Elderly Voice
Speech data from elderly speakers, usually stratified by age band and health status, used for voice products built for older users.
-
Accented English
English speech from non-native speakers or specific regional accents, grouped by accent origin, used to improve accent robustness in ASR.
-
Speech Translation
Parallel data pairing source-language audio with target-language translation, used for speech-to-text and speech-to-speech translation models.
-
In-the-Wild Speech
Speech collected under fully natural conditions — no studio, no topic constraints — as close to real usage as collection gets.
New data paradigms
Embodied AI, synthetic data and world models — categories where the data formats are still being standardized. 3 categories
-
Embodied AI Data
Multimodal data generated as robots and embodied agents operate in real environments — vision, action trajectories, and language instructions.
-
Synthetic Data
Data generated by models or simulation engines, used to fill gaps where real data is scarce, cover long-tail scenarios, and control cost.
-
World Model Data
Environment interaction data for training world models, emphasizing temporal consistency, physical plausibility, and predictability.
Non-speech data
Adjacent categories we source through partner producers. Same sourcing model, different production chains. 15 categories
-
LLM Training Data
Text corpora for LLM pretraining and fine-tuning — instruction data, preference data, and multi-turn dialogue data.
-
Computer Vision Data
Annotated image and video data, with annotation types including classification, bounding boxes, segmentation masks, and keypoints.
-
NLP Data
Annotated text for NLP tasks — named entities, dependency syntax, sentiment polarity, and relation extraction.
-
Multilingual Language Data
Parallel or non-parallel text corpora spanning multiple languages, used for multilingual models and low-resource language research.
-
Translation Data
Parallel corpora pairing source and target language, used for machine translation training and evaluation.
-
Dictionary Data
Structured dictionary entries with definitions, parts of speech, pronunciation, and example sentences, used for dictionary products and lexical understanding tasks.
-
Text Data
General-purpose text corpora from web pages, forums, reviews and similar sources, used for language model pretraining and text classification.
-
Video Data
Annotated video footage, with the annotation needed for tasks such as action recognition, object tracking, and temporal segmentation.
-
Satellite Imagery Data
Remote-sensing satellite or aerial imagery, with multispectral bands and geographic annotation, used for land-cover classification, change detection, and agricultural monitoring.
-
First-Person View Data
Video and sensor data captured from the operator's viewpoint with wearable devices, used for embodied AI and action understanding.
-
Autonomous Driving Data
Road-scene data captured by vehicle sensors — cameras, lidar, and annotation boxes — used to train perception models.
-
Legal Data
Legal filings, case law, and statutory text, used for legal search, contract analysis, and legal question-answering models.
-
Patent Data
Patent applications, claims, and citation data, used for patent search, technology trend analysis, and infringement risk detection.
-
News Data
News articles and reportage with date, source, and topic tags, used for event extraction and temporal analysis.
-
Academic Data
Academic papers, abstracts, and citation data, used for research question answering, literature review, and knowledge graph construction.
Submit a sourcing request
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.