Scoping a second-language batch: what transfers from the first

A second-language batch looks like a repeat of the first. It is a new project with recycled parts, and knowing which parts recycle is the scoping work.

The first language batch ends, it went reasonably well, and the natural next step is to repeat it for the next language. The estimate for batch two is derived from batch one, and the plan is a copy with the language name changed.

That plan is wrong in a specific and predictable way. Some parts of batch one are reusable assets, some are reusable only after adaptation, and some look reusable and are not. The budget error in a second batch almost always comes from putting the third category into the first.

What carries over intact

Reusing the contract structure is the highest-value item here, because the risk terms are language-independent and renegotiating them costs the same effort the second time.

  • The recording protocol and equipment configuration, if the deployment conditions are the same.
  • File formats, naming conventions, and the validation script that checks deliveries.
  • The review and acceptance procedure, including the sampling method and rejection criteria.
  • The contract clauses on rework, re-collection allowances, and change pricing that took a round of negotiation to get right.
  • The project roles. The manager and quality lead who ran batch one know where the process actually breaks, which is knowledge that does not transfer through documentation alone.

What has to be rebuilt from zero

Recruitment does not transfer at all. A new language means a new population, new channels, new screening questions, and a new lead time that has to come from a feasibility check rather than from batch one: recruitment speed in the first language is not evidence about the second.

The guideline does not transfer as a whole. Its language-specific decisions — how fillers are written, how numerals are handled, what the script conventions are — have to be made again for the new language, ideally through the same calibration round the first batch used.

Pronunciation and lexicon assets start empty: reference lexicons, name lists, and normalization tables.

What transfers but will mislead you

Three things from batch one look portable and produce bad estimates when treated that way.

The timeline. Batch one had the duration its recruitment phase gave it, and batch two has a recruitment phase of a different length, in either direction. Reusing the duration without redoing the recruitment estimate imports that error into the whole plan.

The per-unit cost. The second language is likely more expensive per unit if it is less resourced, and cheaper only if the first language was unusually difficult. The direction has to be argued from the new population, not assumed from the old one.

The conventions of the guideline. Rules that batch one found natural may not exist in the second language. An orthography that is fully standardized in one language may be contested in another, and importing the rule wholesale produces instructions that annotators cannot follow consistently.

Run a reuse audit

The scoping method is a table with one row per element of batch one — protocol, guideline, recruitment, tooling, contract, roles, timeline model — and one of three labels per row: reuse as-is, adapt, or rebuild.

The audit forces every label to be argued before the budget is set. Two rows deserve special treatment.

Recruitment gets a feasibility test rather than an estimate: attempt to recruit a small number of speakers, three is enough, through the intended channels, and time how long it takes. That number is the foundation of the schedule, and it is the one part of the plan that cannot be reasoned into existence.

The guideline gets a calibration round, the same as batch one. Budget it.

Re-run the pilot, and reuse the error list

The pilot has to be re-run for the new language even if batch one went perfectly, because the pilot tests whether this language conventions hold, and those conventions are new.

There is one piece of batch one worth importing wholesale: its error history. Every rejection reason from batch one is a hypothesis about batch two, and acceptance sampling in the second batch should check those failure modes first. If the same errors appear, the process has a systemic problem. If they do not, the sampling can move on to language-specific risks.

The second batch is also the moment to fix what was tolerated in the first: the review step that was always rushed, the manifest maintained by hand. Batch two inherits the process, not just the data, and the process is cheaper to improve between batches than during one.

More insights

  • Handling a data batch that failed acceptance

    A failed batch is a decision with three possible outcomes, and the expensive mistake is choosing the wrong one. How to measure the failure, and what to agree beforehand.

  • When to stop collecting data

    More data stops helping before it stops costing. The signals that the curve has flattened, and the stopping rule that makes the decision mechanical.

  • Avoiding scope creep in a data project

    In a corpus project every property is additive, so requests arrive continuously and each one looks small. The fix is a classification step before any pricing.

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com