Planning a multilingual data project: sequence, volume, and consistency

Multilingual projects do not fail inside a language. They fail at the seams: sequencing, volumes set by round numbers, and guidelines that fork silently.

A multilingual project is not one project. It is a set of parallel projects sharing a specification, a budget, and a quality function, and the failure modes live at the points where they touch.

Three decisions determine whether it holds together, and all three are made early: the order the languages run in, the volume assigned to each, and how annotation conventions stay aligned as languages multiply. Each has an obvious answer that is wrong often enough to be worth examining.

Sequence before volume

There are two coherent ways to order the languages, and they address different risks.

Process-first sequencing runs one well-resourced language first, to shake out the pipeline — recruitment channels, tooling, the review loop — while mistakes are still cheap. The risk it addresses is a broken process, and a familiar language makes that visible fast.

Critical-path-first sequencing starts the hardest language immediately, because recruitment lead time for an uncommon population dominates the whole schedule. The risk it addresses is a schedule gated by the slowest input, where starting late is unrecoverable.

Most projects need both, in that order: one pilot language to validate the process, launched in parallel with the first recruitment push for the hardest language, so the slow clock starts on day one.

Set speaker targets per language, not hours

Equal hours across languages feels fair and is usually wrong. The property that makes a language dataset useful is the same one that matters in a single-language project — the number of distinct speakers — and the population available per language is not equal.

For a language with a large, accessible speaker pool, a high speaker target is achievable. For a smaller or geographically concentrated population, the same target may be unreachable, and the honest response is to set the target from a recruitment feasibility check rather than from the budget.

The useful artifact is a per-language sheet with identical fields: target speaker count, minimum hours per speaker, recruitment channel, and a realistic lead time. When those sheets exist side by side, the languages that will slip become visible before the budget is committed.

Localize the guideline, not just the specification

The specification is the document of the buyer. The guideline is the document of the annotators, and it is the one that has to work in every language.

Translation is the starting point, not the finish. Terms of art in annotation — what counts as a filler, whether numerals are normalized, how an unintelligible stretch is marked — do not have stable translations, and a translated guideline that is not tested will be interpreted differently in each language.

The control that catches this is a calibration round per language: three annotators independently annotate the same thirty minutes, agreement is measured, and every disagreement is resolved into a sentence of the guideline. This happens before production annotation starts. The output is a language-specific addendum to the shared guideline, not a fork, because the addendum is owned centrally and versioned with the main document.

What carries across, and what forks

Sorting assets into these three groups early prevents most of the avoidable work.

  • Carries across: recording protocol, equipment configuration, file formats and naming, the review procedure, and the acceptance sampling method.
  • Carries across with review: consent documentation, which has to be checked against each jurisdiction rather than reused unmodified.
  • Forks per language: the guideline addendum, pronunciation and lexicon references, and recruitment channels.
  • Never forks: the change log. One document, one owner, one place where a convention change is recorded, even when the change applies to a single language.

The shared bottleneck

Languages can run in parallel. Review capacity cannot. Quality control is the step that touches every language, and it is usually staffed for one.

The consequences are practical: stagger language starts so their review peaks do not coincide, or budget for review capacity that scales with the number of concurrent languages rather than with the number of hours. The second option is more expensive and more reliable, and the choice between them is worth making explicitly rather than discovering it when three languages deliver in the same week.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com