Resolving disagreements in annotation review: arbitration, voting, and what disagreement is telling you

Most disagreement is a message about the guideline, not a failure by an annotator. How to route it: what to arbitrate, when voting helps, and when it hides the problem.

Sort the disagreement before resolving it

Two annotators label the same item differently, and the instinct is to pick one answer and move on. The disagreement deserves one minute of classification first, because the two kinds have opposite fixes.

Slip disagreement is noise: a misread word, a missed token, a boundary dragged by a careless click. It is random, it is what double annotation finds at its floor, and the remedy is attention rather than editing the document. Ambiguity disagreement is signal: both annotators read the guideline correctly, and the guideline does not decide the case. It is systematic — the same item type will split the same way again tomorrow, on a different annotator.

The tell is repetition. Re-check the pattern on a second sample of the same item type. If it recurs, it is ambiguity, and no amount of re-training will remove it, because the annotators are already following the text they were given.

Adjudicate ambiguity, and write the answer back

The standard structure is a third reader with more experience and more authority than the annotators in dispute. Two design choices make it work.

First, the adjudicator decides from the guideline, not from taste, and when the guideline does not reach the case, the decision becomes a proposed addition to it. Every arbitration therefore produces two artifacts: the label for the item, and the paragraph for the guideline.

Second, apply decisions prospectively by default. Re-labeling everything already annotated each time a case is settled turns the project into a permanent rewrite. A dated decision log, applied forward, keeps the corpus consistent enough while the guideline matures.

  • Who adjudicates: one named senior annotator or the project lead — a person, not a committee.
  • What gets adjudicated: items flagged in review, plus disagreements surfaced by double annotation.
  • What gets recorded: the item, the two readings, the decision, and the rule it establishes.
  • When it ends: a time box per item. An adjudication that runs three rounds is a guideline problem, and the case belongs in a guideline workshop instead.

The traps in majority voting

Voting — three annotators, two agree, the majority wins — is cheap, and it is the wrong instrument for exactly the disagreements that matter most.

The first trap is correlated error. Three annotators trained on the same flawed reading of the guideline will agree with one another and with the error, and the vote confirms the misunderstanding while erasing the one annotator who read the text literally. Majority voting works when errors are independent, and annotator errors are independent only when they are slips.

The second trap is that voting hides the count. The information in a two-to-one split is not the winning label; it is that the item was contested at all. A pipeline that stores only the majority label throws away its own best signal about guideline quality.

Where voting is defensible: large, subjective label spaces with no single right answer, such as preference comparisons or some emotion categories, where the goal is a stable aggregate rather than a per-item truth. Even there, keep the raw votes.

Read the pattern of disagreement

The accumulated list of contested items is the most honest report on a project's health, and it reads in three cuts.

  • By phenomenon: what are the contested items about? Disagreement that clusters on one phenomenon is a missing paragraph, not a training gap.
  • By annotator pair: a pair that disagrees more than everyone else is either a training issue or two people worth listening to. Check whether their reading is the literal one before assuming the former.
  • By time: a rising disagreement rate on stable material is a leading indicator of drift, and it usually shows up before any accuracy number moves.

Three habits that make disagreement worse

Some responses to conflict reliably lower quality, and all three are common.

  • Splitting the difference — forcing a middle label onto an item that needs a decision produces a corpus that matches no annotator and no written rule.
  • Resolving disputes in chat threads, where the decision is real but unrecorded, and the next annotator to meet the case resolves it differently.
  • Treating a high disagreement rate as a personnel problem before checking the guideline. The document is the more common cause, and the cheaper one to fix.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com