Detecting drift in a long annotation project before the batch is ruined

Quality does not fall off a cliff; it slides. The few measurements that catch the slide, and how to tell annotator fatigue apart from a guideline that has loosened.

Drift is a slope, not an incident

A bad batch announces itself. Drift does not. Each week is a small, defensible step from the week before, and the material coming out at month six is not what the guideline produced in week one, with no single decision anyone can point to. By the time a spot check makes it visible, a large share of the corpus has already been produced under the drifted convention, and the cost of the discovery is the cost of re-doing it.

The defense is not more inspection. It is a small number of stable measurements, repeated on a fixed schedule, so the slope is visible while it is still a slope. Three instruments cover most of it: a frozen gold sample, a double-annotated slice, and one distribution count.

The measurements worth repeating

Keep the set small enough that it actually runs every period, and freeze the definitions. A metric whose definition changes is not a trend line; when the gold sample is replaced, re-baseline everything on the same date and mark that date on the chart.

  • Gold score, per week or per batch: accuracy against a fixed, verified sample. This is the primary drift instrument, and the others exist to explain it.
  • Double-annotation agreement on fresh material: two annotators, one slice, a fixed agreement metric. It catches the case where the team converges on a shared new convention — the gold score stays flat because everyone moved together, and only agreement on new items shows the move.
  • A label distribution count: the share of each mark in a rolling window — unintelligible marks per thousand tokens, a particular tag category per hour. Distribution shifts are usually where drift becomes visible first.
  • Throughput, as a supporting number: items per hour per annotator. It does not measure quality, but a step change in it is a reason to look.

Fatigue and a loosened guideline look different

Both show up as a falling gold score, and they need different fixes, so the first job is telling them apart.

Fatigue is within-person and within-day. Quality sags late in a shift, in the last hour before a deadline, on the longest files. Its error signature is slips — dropped words, typos, boundary misclicks — errors the annotator would reject immediately if shown their own work. It concentrates in individuals and in hours, and it recovers after rest or a change of task.

A loosened guideline is between-people and stable in time: a convention that has quietly shifted for everyone. Filler words now being dropped, "unintelligible" now used for merely difficult words, punctuation added where the guideline said none. Its signature is consistent, defensible-looking output that differs from the old convention — the annotators are not making mistakes by their own lights, and re-reading the document is what exposes the gap between what it says and what the floor does.

The test that separates the two: pull ten items from the beginning of the project and ten from the present, and re-check all twenty against the guideline as written. Fatigue shows as scattered errors. A loosened convention shows as a pattern — the same decision, made differently.

Turnover is the quiet multiplier

Long projects lose people. When half the annotators on a batch were not in the room when the guideline was briefed, the guideline is no longer the shared document — it is a document plus folklore, and the folklore wins on the cases that matter.

Two practices keep this bounded: a short re-briefing on a fixed schedule that walks through the current version and the decision log, and onboarding through the gold set, so nobody annotates deliverable material before passing the gate. A project that makes the gold set the entry requirement has a much flatter drift curve, because every new annotator starts calibrated instead of converging over their first two weeks.

Correct without re-doing everything

When drift is found, the instinct is to re-annotate the affected period. That is the most expensive response available, and it is usually the wrong first move.

  • Quantify the window first: find the earliest measurement that shows the shift and count the affected material. The window is often narrower than it feels.
  • Fix the cause before the material: re-brief the team, write the decisions that were being made informally into the guideline, and re-baseline the metrics.
  • Repair by stratum where the drift was mechanical — a tag applied too liberally can be corrected by targeted re-review rather than full re-annotation.
  • Accept and disclose what is not worth repairing. A mild convention change across one early month is often better documented in the dataset notes than re-done, because re-annotation creates a new vintage with its own problems.

Then re-baseline, and say so

After any correction, the old trend line no longer describes the data, and pretending it does is how a project loses track of its own history. Record the date, the cause, and the fix alongside the charts, so that the next person to look at the numbers reads a slope and its explanation rather than an unexplained step.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com