Outsourcing data annotation: what the handover really involves
Outsourcing moves the annotation work, not the responsibility. What you hand over, what the provider must be able to decide alone, and what should stay with your own team.
What outsourcing actually transfers
Sending files to an annotation provider looks like a simple transaction: you have raw material, they have annotators. What actually changes hands is a set of decisions. Before production starts, every convention the annotators will apply has to come from somewhere — your guideline, their habit, or a series of individual guesses.
That is why annotation outsourcing fails in two opposite directions. Under-specify, and the provider invents conventions you discover at delivery, after the whole batch has been labeled the same wrong way. Over-specify, and every ambiguous item becomes a question back to you, the queue grows, and the provider cannot hold the throughput the price assumed.
The work of outsourcing well is drawing the line between those two, and the line has to be drawn before the first file is annotated rather than after the first delivery is rejected.
What you hand over, and in what state
The handover is a small set of documents and access grants. Each one exists to remove a class of questions from the annotator queue:
- The annotation guideline as a versioned document, including every convention decision: filler words, numerals, casing, how an unintelligible stretch is marked, and how a speaker correcting themselves mid-sentence is written.
- The label taxonomy with boundary definitions. A label list without boundary cases is not a taxonomy, because the cases that matter are always the ones sitting between two labels.
- Reference examples: a handful of accepted items per stratum, each with a sentence on why it is correct. These become the calibration material the provider works from internally.
- The escalation path in writing: who the provider asks, through which channel, and what response time they should expect.
- The data itself with a manifest, so both sides can reconcile what was sent against what was received before any annotation is billed.
The intake step most buyers skip
Before production annotation begins, the provider annotates a small sample and sends it back for review. This is not a pilot in the project sense. It is a calibration round, and its purpose is to test whether the guideline is being read the way you meant it.
Score the sample against your own reference. Where the provider differs, the difference is usually not carelessness. It is a sentence in the guideline doing less work than you assumed, and the fix is an edit to the document rather than a complaint to the provider.
Keep the calibration material. The items annotated during intake, with your corrections attached, become the first part of the reference set used for acceptance sampling later, and keeping them costs nothing extra.
What the provider has to decide without you
A provider that must ask about every ambiguous item is slow, and the slowness is structural rather than a staffing problem. The useful split is a rule for the common case and a channel for everything else.
- Give one decision rule that covers most ambiguous items, stated in a sentence an annotator can apply without judgment. If no such rule exists, say so and route the whole category to escalation instead of pretending otherwise.
- Name the adjudicator. One person on your side owns edge-case decisions, and that person has to be reachable during the annotation window, not only at delivery.
- Set a loose question budget: a rate of escalations that is normal, and a rate high enough to signal the guideline has a gap. A provider that escalates nothing is not resolving ambiguity; it is hiding it.
- Ask for a periodic escalation summary. The list of questions asked in a week is the most honest description of your own guideline you will ever receive.
What should not leave your team
Four things stay on your side of the line, and the reason is the same for all of them: they are the definitions of correct, and a definition that lives with the party being measured stops being a check.
- Ownership of the guideline. The provider can propose changes, and the version that applies is the one you published.
- The answer key. If the provider holds the full reference set with your corrections on it, they can annotate to the test rather than to the guideline. Share examples freely and keep the acceptance reference set.
- Final acceptance. The provider can and should run internal quality control. The decision to accept a delivery is yours, made against a sample that you or your reviewer drew.
- The definition of what correct means for your model. Whether a particular edge case matters is a product question, and no provider can answer it from the data alone.
The rhythm after the handover
Once production starts, three things keep the arrangement honest, and all three are cheap: a weekly agreement sample where you re-annotate a small random draw and compare; a standing review of the escalation list; and a change log for the guideline, because the guideline will change and an unversioned change leaves a seam in the dataset.
At the end of the engagement, ask for the provider's error taxonomy, meaning their own list of what was hard and why. It is the most useful artifact a project produces and the one nobody thinks to request. It tells you what the next batch in the same language or condition should watch for, whether or not the same provider runs it.