Buying data for low-resource languages: what changes

For a language with no existing corpus, everything about the project is different — starting with the fact that the timeline is set by speaker recruitment, not recording.

Low-resource does not mean few speakers. Navajo has fewer than 200,000 speakers, but so do several languages with well-developed digital corpora. What makes a language low-resource is the absence of existing text and audio that a model can learn from.

For a buyer, that changes the shape of the project more than any other single factor.

The bottleneck moves

In a well-resourced language, the constraint is usually studio time and annotation throughput. Speakers are easy to find and the recruitment step is a formality.

In a low-resource language the constraint is almost always recruitment. Qualified native speakers may be geographically concentrated, may not be reachable through standard channels, and may have limited availability. Finding forty of them can take longer than recording all of them.

This is why timelines for low-resource projects are unreliable in a specific way: the estimate is dominated by a step whose duration is hard to predict in advance.

Existing text is not a shortcut

When a language has some written material — religious texts, colonial-era documentation, community publications — there is a temptation to use it as the transcription target.

It usually does not work. Older written material tends to be in a register nobody speaks, in an orthography that has since changed, or in a dialect other than the one you are recording. The audio and the text end up describing different things.

What to expect on quality

The realistic quality target for a first collection in a low-resource language is lower than for a well-resourced one, and that is a property of the situation rather than of the producer.

  • Annotator agreement will be lower, because the orthography may not be fully standardized and there is no established convention to fall back on.
  • Some of the annotation budget will go to establishing conventions rather than applying them.
  • A pilot batch matters more here than anywhere else, because it is where the conventions get tested.

The offsetting advantage

Low-resource data is scarce, which means it is also defensible. A competitor cannot buy the same corpus from a marketplace, because there is no marketplace listing for it.

For teams building a capability that depends on language coverage, that scarcity is the point. The project is slower and more expensive per hour, and the resulting asset is one that cannot be replicated by writing a larger purchase order.

More insights

  • Sourcing speakers of a rare language

    The constraint is not finding people who speak it. It is finding people who meet every other requirement at the same time, and proving that they do.

  • Managing a multi-language data program

    Quality means something different in every language, and one delayed language can hold a whole release. Both problems are managed with the same three artifacts.

  • Handling a data batch that failed acceptance

    A failed batch is a decision with three possible outcomes, and the expensive mistake is choosing the wrong one. How to measure the failure, and what to agree beforehand.

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com