Data labeling vendor selection criteria you can score and defend
A weighted scorecard with defined anchors, a disqualification floor, and a scoring order that keeps price from quietly deciding the ranking.
A criteria list is not a scorecard
Most selection criteria lists fail for the same reason: they name the dimensions without saying what a high score looks like or how much each one matters. The result is a ranking that reproduces the buyer's first impression with more steps and more confidence.
Three things turn a list into an instrument. Anchors, so a score means the same thing to two reviewers. Weights, fixed before any proposal is opened. And a floor, because a weighted average can hide one fatal weakness behind eight good scores.
The criteria and their weights
The weights below are a starting point rather than a standard, and they total one hundred. What matters is that they are written down and argued before the responses arrive.
- Specification comprehension and pushback — 15. Did the provider engage with the brief, flag conflicts, and ask what the data is for? Or agree with everything?
- Guideline and calibration process — 15. Does a written guideline exist, is there a calibration round before production, and who owns the document?
- Measured quality process — 15. What is measured, on what sample, by whom, and is the measuring function independent of the producing function?
- Reach into your population — 10. Recruitment channels for the languages, regions, and speaker profiles you need, with a lead-time estimate you can check.
- Throughput and concurrent capacity — 10. How many projects run at once, what the current queue looks like, and what happens to your batch when a larger client arrives.
- Data handling and confidentiality controls — 10. Access rules, retention clocks, deletion verification, and how material on annotator machines is handled.
- Contract and risk terms — 10. Rework allowance, re-collection terms, change pricing, cure period, and which side carries which risk.
- Communication and escalation behavior — 5. Named contacts, response substance, and how the provider behaved during the procurement itself.
- Price on a normalized unit — 10. Cost per accepted unit as your acceptance procedure defines it, not cost per quoted unit.
Anchors: what the numbers mean
An anchor is a sentence describing the evidence a score requires. Without them, two reviewers land two points apart on the same proposal and the reconciliation turns into an argument about taste.
Worked examples, for the two criteria carrying the most weight:
- Guideline, score 5: a versioned guideline was shared without being asked for, and the provider described a calibration round with a measured agreement figure. Score 3: a guideline exists and is shared on request, with no calibration step described. Score 1: conventions are described as standard practice with no document behind the claim.
- Quality process, score 5: a named reviewer independent of production, a stated sample fraction, and a stated metric with a reference set. Score 3: internal review described, but the reviewer also produces the work. Score 1: quality described in adjectives, with no sampling rule and no metric.
- Price: score against cost per accepted unit, and state in the anchor which costs are included. A quote with an open-ended change process scores lower than its unit figure suggests.
The disqualification floor
Some findings end the evaluation regardless of the weighted total. A weighted average is built to absorb a weak area, which is exactly wrong when the weak area is a missing control rather than a thinner capability. Check these before the weighted scoring starts, because removing a candidate on a missing document is faster than scoring them, and a weighted average would otherwise let a control gap survive into the ranking.
- No written guideline, or a refusal to share it.
- Cannot name who performs the annotation, or refuses to list subcontractors by category.
- No acceptance procedure, and no reference set to measure quality against.
- No described path for deleting material when the project ends.
- Refusal to run a paid pilot, or an offer of a free pilot produced by the sales side.
The scoring order that keeps it honest
Procedure matters more than the instrument, because the common failure is a scorecard quietly reverse-engineered from the price list. Five rules prevent it:
- Fix the weights and anchors before the proposals arrive, and treat any later change as a documented decision with a stated reason.
- Score the technical criteria with the price pages removed from the document.
- Require evidence for every score: which answer, which attachment, which page. A score with no citation is an impression.
- Have two people score independently, then reconcile in writing while keeping both original sheets.
- Keep the same criteria across procurements. A scorecard rewritten each time cannot compare candidates over time or reveal that your own requirements are drifting.
Where a scorecard stops being useful
Two honest limits. First, when only one plausible provider exists there is no selection to make, and the negotiation becomes one about scope, schedule, and risk terms. Running a weighted scorecard against a single candidate produces a number that only pretends to be a decision.
Second, the scorecard chooses who gets a paid pilot, not who gets the production contract. The pilot is where quality claims meet your actual material, and it can overturn the ranking. What the scorecard buys you is a short pilot list and the written reasons behind it.