Writing an annotation guideline that holds up under production

The guideline is what makes two annotators interchangeable. What belongs in it, how to write boundary cases, when examples earn their place, and how it changes.

The guideline is the product, not the paperwork

Ask what makes two annotators interchangeable and the answer is never that both are careful. It is that both are working from a document with an answer for the cases they will meet, including the cases nobody thought of when the project was scoped. The guideline is what converts individual judgment into a repeatable process, and a repeatable process is what a buyer is actually paying for.

The common failure is a guideline written as principles. "Transcribe verbatim" and "mark unclear speech" are principles, and they settle none of the twenty questions an annotator meets in the first hour: is "gonna" written out; does a cough get a token; is a false start kept; is a name spelled as heard or as written; does background speech from a second person enter the transcript.

Give every phenomenon a rule, then an example

The working unit of a usable guideline is a phenomenon with a rule attached: what it is, how to mark it, and what not to do. The phenomenon list in transcription work runs long — fillers, false starts, repetitions, numbers, dates, loanwords, names, acronyms, unintelligible stretches, background speech, non-speech sounds, mid-sentence language switches.

Examples earn their place when they carry the reasoning. An example without a reason teaches pattern matching, and pattern matching breaks on the first case that is close to the example but not identical. The format that holds up: two or three examples per rule — one clear case, one clear counter-case, and one borderline case with a sentence explaining which way it falls and why. The borderline example is the one annotators will actually consult at the moment of doubt.

  • Write the rule as a decision, not a value: "if the speaker restarts within two words, keep both attempts and mark the restart" rather than "be faithful to what was said."
  • Show the counter-example so the rule's edge is visible, not just its center.
  • Write the reasoning on the borderline case, in one sentence, in the guideline itself.
  • Say what to do when a case is not covered: flag it and continue, rather than decide alone and silently.

Length is a symptom, not a target

Guidelines are usually too short, not too long, but length is the wrong thing to manage. The right question is coverage: does the document have an answer for the cases that will actually appear? A practical test is to collect twenty hard cases — from a pilot, from a sample of real recordings, from the annotators themselves — and check each one against the guideline before production starts. Every case with no answer is a paragraph that will otherwise be written by whoever hits it first, inside their own head, and applied only to their own work.

After coverage, structure decides whether the document gets used. A phenomenon index, one section per phenomenon, decisions numbered so they can be cited in review, and a visible version number on every page. A reviewer who can cite rule 4.2 resolves a dispute in minutes; a reviewer arguing from taste takes rounds.

The change process is part of the guideline

A guideline that cannot change will be abandoned the first time it is wrong, and a guideline that changes silently makes earlier work incomparable. Both problems have the same fix: versioned changes, a change log, and a stated rule for how each change applies.

  • Number every version and date every entry in the log; "the current guideline" is not a version.
  • Say whether a change applies forward only or triggers rework of already-annotated material. The default should be forward only, with rework reserved for changes that flip labels rather than refine them.
  • Record why each change was made — the case that forced it — so the reasoning survives the people who were in the room.
  • Announce changes in a briefing rather than by attachment; the point is that everyone reads the new paragraph on the same day, the same way.

Test it on people who were not in the room

Before production, have two people annotate the same hard cases using only the guideline, with no discussion. Then compare. The disagreements are not a problem with the annotators; they are the list of paragraphs that still need writing, and at this stage they cost an afternoon to fix.

Run the same test again on a small sample mid-project, and it answers a different question: is this still the document governing the work? If two annotators who were both trained on the current version disagree as though they had never seen it, the guideline has been replaced in practice by whatever the floor believes. That is the moment to re-brief, not the moment the delivery is rejected.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com