What a paid pilot batch should prove before you commit to the full project
A pilot is not a free sample. It is a test with pass and fail conditions written in advance, and it should prove four things a proposal cannot.
A pilot batch exists to answer one question: can this producer deliver what the specification actually requires, at the quality level and the rate the full project depends on. Everything else about a pilot is secondary.
That framing matters because pilots are often treated as samples, judged on how the finished material looks. A sample is evaluated on the best material in it. A test is evaluated on the worst, because the worst accepted unit is what defines the floor of the dataset you are buying.
The four things a pilot has to prove
Ask for the pilot to be delivered with the same artifacts as the full project: the guideline, the speaker manifest, the quality report. If those artifacts do not exist at pilot scale, they will not exist at production scale.
- Specification comprehension: that the producer read the specification the way you meant it, including the parts that were ambiguous.
- The quality ceiling: how good their best work gets, which sets the expectation for the whole batch.
- The quality floor: how bad the worst accepted unit is, because that is what the model actually trains on.
- Throughput: how long a defined batch takes end to end, including the review round, which is where timelines usually slip.
How large is large enough
A pilot that is too small tests nothing except the ability to pick good examples. Two sizing rules help.
First, the pilot has to cover every stratum you specified. If the project includes four age bands and three recording environments, the pilot needs at least one example of each combination, or the review cannot tell you whether all of them are handled.
Second, the pilot has to include one complete review cycle: delivery, your feedback, their rework, your acceptance. A pilot that ends at first delivery proves the producer can produce. It does not prove they can revise, and revision is where most working relationships fail.
Write the pass and fail criteria before the pilot starts
The criteria have to exist before any material is produced, or they will be reverse-engineered from whatever arrives.
Where quality can be measured, measure it: word error rate on a held-out sample, agreement on a double-annotated subset, format compliance checked by a script rather than by eye. Where it cannot be measured, such as audio that is technically correct but wrong for the use case, say who decides and on what basis, and make that person available during the review window.
One practical rule: define what fraction of the pilot will be inspected. Reviewing everything is not scalable, and reviewing only the best-sounding files is not a review. A random sample plus every file the producer flagged as borderline is a defensible middle.
What should stop the project
A pilot is the cheapest place to end a procurement, and these are the signals that justify it.
- Two consecutive review rounds fail on the same criterion. The first failure is information; the second is a capability limit.
- The pilot itself slips its date. Pilots are small and controlled, so a broken schedule at this size will be worse at production size.
- Rework introduces new errors in material that previously passed, which means the review process is not under control.
- The producer pushes back on written acceptance criteria and wants to be judged on overall impression.
- The pilot was produced by a different team than the one that will run production.
The pilot that passes and still proves nothing
A pilot can be run honestly and still tell you the wrong thing. The usual causes: speakers cherry-picked rather than sampled the way production will sample; a reviewer who was the producer rather than an independent check; and a batch produced at a pace that the full project makes impossible.
The defense is to specify how the pilot participants are selected, to insist that the pilot is paid on the same rate structure as production, and to treat any offer of a free pilot produced by the sales side as a preview rather than a test. Free samples are marketing. A paid pilot is evidence.