What is Translation Data?
Parallel corpora pairing source and target language, used for machine translation training and evaluation.
What buyers get wrong
Quality tiers within parallel corpora vary enormously — human translation, post-edited MT, and raw MT mixed together drag the whole set down, so they have to be delivered separately.
The specification detail that decides everything
Translation method (human / post-edited / machine) must be annotated per segment, not described as a blanket note.
This is the line item that most often separates a dataset that works from one that gets re-ordered. It belongs in the specification before production starts, not in the delivery review.
How we source translation data
We do not hold inventory in this category. A requirement comes in, we match it to a producer who can meet the specification, and we agree the quality bar with both sides before recording begins. A pilot batch goes first so misalignments surface early.
Specification at a glance
| Category | Translation Data |
|---|---|
| Group | Non-speech data |
| Critical spec | Translation method (human / post-edited / machine) must be annotated per segment, not described as a blanket note. |
| Included by default | Pilot batch, metadata schema, source and consent documentation |
| Pricing | Quoted from specification — no published rates |
Related categories
-
LLM Training Data
Text corpora for LLM pretraining and fine-tuning — instruction data, preference data, and multi-turn dialogue data.
-
Computer Vision Data
Annotated image and video data, with annotation types including classification, bounding boxes, segmentation masks, and keypoints.
-
NLP Data
Annotated text for NLP tasks — named entities, dependency syntax, sentiment polarity, and relation extraction.
-
Multilingual Language Data
Parallel or non-parallel text corpora spanning multiple languages, used for multilingual models and low-resource language research.
-
Dictionary Data
Structured dictionary entries with definitions, parts of speech, pronunciation, and example sentences, used for dictionary products and lexical understanding tasks.
-
Text Data
General-purpose text corpora from web pages, forums, reviews and similar sources, used for language model pretraining and text classification.
Questions about translation data
Do you have translation data available now?
No — we do not hold inventory. Everything is produced against a specification. That means the timeline starts when the spec is agreed rather than immediately, and it also means the data matches your requirement instead of approximately matching something already built.
Can this be combined with other categories in one delivery?
Yes, and it is common. A single specification can cover several categories across several languages, with one pilot batch and one quality bar. It usually reduces total cost because speaker recruitment and project setup are shared.
What does the consent documentation cover?
Signed speaker consent stating the intended use, the collection methodology, and the annotation guideline the data was produced under. For projects involving EU data subjects the chain is built to satisfy GDPR, including the transfer mechanism.
Request translation data
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.