Building a gold set for annotation QC: composition, leakage, and when to retire items

A gold set is how you measure annotators instead of trusting them. How to build one that keeps measuring something, and the ways it quietly stops working.

What the gold set is for

A gold set is a small collection of items whose labels are settled — reviewed, argued over, and recorded as correct — fed back through the annotation pipeline as if it were ordinary work. Comparing what comes out against the gold labels measures the one thing a review process cannot measure about itself: whether the annotators, as a group, are still calibrated.

That makes it a different instrument from a model test set. A test set measures a system against fixed truth once. A gold set measures people against fixed truth repeatedly, and the useful output is the trend across weeks, not any single score.

Composition: boundary-heavy, label-complete, small

A gold set of 100 to 300 items covers the usual purposes. Below about 50, per-item noise dominates and the trend line stops being readable.

The composition rules that keep it informative:

  • Cover every label and every phenomenon the guideline names, a handful of items each, so a defect in one area cannot hide inside a good average.
  • Overweight boundary cases on purpose — the cases the guideline had to argue about. Clean, easy items tell you almost nothing, because everyone gets them right.
  • Where the honest answer is that two readings are both acceptable, record that status and score those items separately. A gold set that pretends every case has one right answer ends up punishing correct judgment.
  • Keep the difficulty mix stable across versions, or a score change will be indistinguishable from a composition change.

Who labels them, and why the reasoning matters

The people whose work the set will score cannot be the people who decide its answers. The workable chain is a senior annotator as a first pass, a second senior annotator working blind to the first, then adjudication of the differences with the reasoning written down. Every gold item should be able to answer "why" in one sentence, because that sentence is what a reviewer uses when an annotator challenges a score.

One practical detail: label the gold items at the same granularity as production. A gold set of clean single utterances cannot score a pipeline that also handles overlapping speech and long multi-speaker files.

Leakage: the failure that makes a gold set decorative

A gold set stops measuring the moment it becomes memorizable. In practice it leaks in one of four ways.

  • Recognizable items: files or ids the annotators have already seen in training material, or items that look different from the surrounding batch — a different format, a different naming pattern, suspiciously clean audio.
  • Position: gold items clustered at the start of a batch, or always arriving in the same order.
  • Feedback: publishing the answer key after each round, so annotators optimize against the key instead of the guideline.
  • Telling: an interface or a batch note that signals which items are checks. The annotator now knows exactly when to slow down.

Run it the way production runs

The countermeasures are the mirror image of the leaks. Draw gold items from the same distribution as production work, embed them at a low rate among ordinary items, shuffle their positions, and never mark them in the interface. Keep the answer key outside the working system, and rotate items out on a schedule: an item everyone now passes has stopped carrying information, whatever its history.

One boundary worth stating: the gold set is not a training corpus. Using gold items as worked examples, then scoring on the same items, measures recall of the answer key. Keep the examples used in training disjoint from the items used to score, or accept that the score means nothing.

Read the scores correctly, and maintain the set like a document

A gold score is a screening instrument, not a verdict. Read per annotator, it is the cheapest gate before a new person touches deliverable data. Read per item, a strong annotator missing an item usually means the item is bad — the label was wrong, the case is genuinely ambiguous, or the audio is unusable. Read per batch, the group average over time is the drift signal; individual variation is normal, and the group moving is not.

Maintain the set the way the guideline is maintained: versioned, with a change log of items added, retired, and corrected. When the guideline changes a decision, the gold items it affects get re-labeled in the same change, because a gold set running on an old guideline punishes annotators for following the new one. A quarterly review, a saturation rule for retiring items, and an intake rule for new phenomena are enough to keep it alive.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com