Avoiding scope creep in a data project

In a corpus project every property is additive, so requests arrive continuously and each one looks small. The fix is a classification step before any pricing.

Why data projects attract scope creep

A data project has no natural definition of done. Software has a release, a building has a handover, and a corpus has whatever property the buyer thought of last. Every dimension of a dataset is additive — more speakers, another recording condition, deeper annotation, an extra language — and adding to it always sounds reasonable.

The second reason is visibility. The buyer can see material as it is produced, in a pilot or a first batch, and real data generates requests that no specification meeting would have produced. This is not a failure of discipline on either side. It is a property of working with a deliverable that can always be made slightly better.

Classify before pricing

The usual mistake is to price a change request on its size. Size is the wrong first question, because two requests of identical size can differ by an order of magnitude in cost depending on whether the material they touch has already been produced.

Sort every request into one of four classes, and let the class decide who absorbs it.

  • Class one touches material that does not exist yet. A wording change in the prompt set, an extra participant in a stratum that has not started. Cheap, and normally absorbed.
  • Class two invalidates material already produced. A change to the segmentation rule, a new requirement about background speech, a corrected pronunciation. Expensive, because it carries the rework plus the seam between the old rule and the new one.
  • Class three changes the process for everything that follows. A guideline revision, an added recording condition, a switch of annotation tool. This is a versioned decision rather than a request, and it needs a named owner on both sides.
  • Class four is not a change at all. More hours, another language, a second modality. It is a new order and should be quoted as one, even when it arrives inside a sentence that begins with while you are at it.

The seam is the real cost

When a rule changes mid-production, the corpus holds material annotated under two different rule sets. The two are not interchangeable, and a model trained on the mixture learns both, which is usually not what anyone wanted.

Three ways to handle a seam, in order of preference. Apply the new rule from the next batch onward and record the boundary in the manifest, which is cheap and leaves a corpus with two documented eras. Re-annotate the affected slice so the corpus is uniform, which is expensive and clean. Or drop the affected slice, which is sometimes the cheapest honest option when the slice is small.

What must not happen is an undocumented seam. A corpus with two rule sets and no record of where they meet produces evaluation numbers nobody can explain later, and by then the explanation is usually unavailable.

What to accept, and what to refuse

What follows is a decision rule rather than a policy, and the value is in applying it the same way every time. An inconsistent answer teaches the producer that asking again is worth the effort.

  • Accept a change that fixes a genuine ambiguity in your own specification. If two readings were possible and the producer took the other one, the cost is yours, and absorbing it once is cheaper than the argument.
  • Accept small additive requests that invalidate nothing, when the schedule has room. The goodwill is real and the cost is bounded.
  • Refuse, or move to a new order, anything that changes what the dataset is for. A corpus that was going to train a recognizer and is now also going to evaluate speaker verification is two datasets, and the second one has requirements the first never had.
  • Refuse late format changes with no functional benefit. Repackaging a delivery in the last week risks the parts that already passed.
  • Defer anything that can be added later without touching existing material. Extra metadata is usually deferrable; changed annotation is not.

The change order, and who signs it

One page, numbered, appended to the contract. It records the request, the date, the requester, the class, the effect on cost, the effect on schedule, the effect on consistency, the decision, and two signatures. A request without a change order did not happen, and the rule has to hold in both directions, including for changes the producer wants to make.

Name one approver per side, and give each a defined allowance they can spend without paperwork, expressed in hours of work rather than in currency. Without an allowance, trivial requests consume the process. With an allowance that is too large, class two changes get absorbed by accident and nobody notices until the seam appears.

Track the change rate as a number rather than as a feeling. Change orders per week, and rework hours as a fraction of production hours, are both cheap to record and both move before the schedule does.

Re-baseline instead of accumulating

Small changes do not stay small. When the cumulative effect of accepted changes crosses a fraction of the original scope — a fifth is a reasonable line to draw — the plan should be re-baselined once, formally, rather than adjusted a sixth time.

Re-baselining means re-issuing the scope, the schedule, and the acceptance criteria as a new version with its own date, and closing the earlier version. It gives both sides one clean reference point and stops the slow drift in which every party is working from a slightly different memory of what was agreed.

The alternative is the project that ends with everyone agreeing the work was reasonable and nobody agreeing it was what they ordered.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com