How to estimate a data collection timeline: two clocks and one critical path

Most timeline estimates add up work that runs in parallel and ignore the phases that wait on other people. Separating the two clocks fixes both errors.

Two clocks, and why adding them up fails

A collection timeline estimate usually goes wrong for one of two reasons. Either it treats waiting time and working time as the same kind of quantity, or it assumes that steps which depend on each other can overlap. Both errors come from writing the plan as one list of durations and summing it.

Split the project into two clocks instead. The throughput clock measures work: hours of audio divided by hours recorded per day, annotation hours divided by annotator-days, inspection minutes per sampled unit. It shrinks when capacity is added. The lead-time clock measures waiting: recruitment, legal review, equipment delivery, a participant who can only record on Saturdays. It does not shrink when capacity is added, because the delay is not yours to spend.

Most projects end up with a throughput clock measured in weeks and a lead-time clock measured in months, and the lead-time clock decides the delivery date.

Place the lead-time items first

Start with the phases where the project is waiting on someone outside it. These have the least flexibility and the widest uncertainty, and they need to be on the calendar before any rate is computed.

  • Consent and legal review. External counsel works to its own schedule, and a new jurisdiction or a new population adds a review that did not exist on the last project. Estimate in weeks, and start it before the production plan is finished.
  • Recruitment to target. Not the time to find one speaker, but the time to find, screen, and schedule the whole pool. This is the most uncertain number in most projects, and the honest way to get it is a probe: try to recruit three people through the intended channels, and time how long that takes.
  • Site or equipment access, when the project needs either. A rented room, a portable recorder, a booth booking, or permission to record in a community space all sit in their own queue.
  • Per-participant consent, which is bounded by the participant's availability rather than by the project's.
  • Review rounds that involve a third party, such as a client legal team or an ethics board. These meet on a fixed cadence, so a two-day review can take two weeks to schedule.

Then the throughput items

Only once the waiting phases are placed does it make sense to compute rates. A rate is always per day per unit of capacity, and it should be written that way rather than as a total.

Recording: hours per session, sessions per day, and how many sessions run at once. Parallel sessions are limited by rooms, recorders, and session leads, not by speakers alone. A plan that assumes six simultaneous sessions while the project owns three recorders has a hidden dependency in it.

Annotation: hours per annotator per day, the number of annotators, and a separate review pass. Budget the review as its own rate rather than as a percentage, because review capacity is usually the thinner resource and the one that gets borrowed from.

Inspection: the sampling rate multiplied by the time to inspect one sampled unit. If acceptance examines a tenth of the output, the inspection rate is what determines the gap between production finishing and delivery, and that gap is often a week or more.

Delivery preparation: manifest generation, checksums, a validation run, packaging. Small in hours, but it is serial and it lands at the end of the project when everyone is tired.

The critical path is shorter than the plan

Once both clocks are laid out, the estimate is the longest dependent chain rather than the sum. Several of the phases above run alongside each other, and the point of the layout is seeing which ones actually touch the delivery date.

On a first project in a new population, the critical path usually runs through recruitment, then the pilot, then the pilot review, then production. The pilot sits on the critical path even though it produces a tiny share of the data, because the guideline is not settled until the pilot review closes and production cannot start before that.

Two consequences follow. Shortening production does not move the date when recruitment is the constraint. And the pilot review needs to be a dated milestone with named participants, because a review that waits for people to become available is where a week disappears without appearing in any estimate.

What cannot be compressed, and what can

Compression is available on one side of the ledger only, and forcing it on the other side converts a schedule problem into a quality problem.

  • Cannot be compressed: the lead-time items, the pilot review round, the calibration round for a new language or a new guideline, and the first session with a new speaker or new equipment.
  • Can be compressed, at a cost: recording throughput by adding rooms and session leads; annotation throughput by adding annotators, up to the point where the guideline becomes ambiguous and disagreement rises; inspection by sampling instead of reviewing everything, which trades detection probability for calendar time.
  • Looks compressible and is not: the final week. Moving delivery preparation earlier does not help, because the manifest and the validation depend on the last accepted file.
  • Compressible only by repeating work: a re-session. Adding capacity to replace a failed batch costs more per unit than the original production, because one speaker's session cannot be batched with anybody else's.

How much buffer, and where it goes

Buffer belongs at the end of the critical path as a visible line, not distributed invisibly across every phase. Distributed buffer is why a plan slips while each individual phase reports as on track.

Size it by novelty rather than by habit. A repeat project in a familiar language, with the same producer and the same equipment, is predictable and needs a modest allowance. A first project in a new population or a new country is not, and almost all of the uncertainty sits in the lead-time items rather than in the production rates.

The practical output is three dates rather than one: the date the plan produces if everything runs at the estimated rate, an earlier date that stays credible if the waiting phases come in fast, and a later date that accounts for one re-session and one extra review round. Planning against a single date hides the fact that the distribution is lopsided, and the lopsidedness is the useful part.

Finally, name the two moments when the estimate gets revised: the end of recruitment and the close of the pilot review. Both arrive early enough to act on, which is what makes them worth marking.

More insights

  • Handling a data batch that failed acceptance

    A failed batch is a decision with three possible outcomes, and the expensive mistake is choosing the wrong one. How to measure the failure, and what to agree beforehand.

  • When to stop collecting data

    More data stops helping before it stops costing. The signals that the curve has flattened, and the stopping rule that makes the decision mechanical.

  • Avoiding scope creep in a data project

    In a corpus project every property is additive, so requests arrive continuously and each one looks small. The fix is a classification step before any pricing.

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com