How to compare quotes for a data collection project fairly
Two quotes for the same project can differ by a factor and still describe identical work. The gap is usually the denominator, and normalizing it is the job.
Quotes for data collection are not comparable by default, and the reason is arithmetic rather than dishonesty. Each producer prices against its own reading of the scope, using its own unit, with its own assumptions about what is included. Put the numbers side by side and the comparison is meaningless in a specific, measurable way.
The fix is not to demand a lower number. It is to force the quotes onto a common basis before reading them, which is mostly preparation on your side of the table.
Normalize the denominator first
The unit in a quote is rarely the unit you think it is. A per-hour figure may refer to recorded audio, delivered audio, accepted audio, or annotation hours on top of recording. Two quotes that look far apart can be identical once the denominators are aligned, and the reverse happens too.
Pick one unit — accepted delivered hours, defined by your acceptance procedure — and restate every quote in it. Where a producer cannot restate the quote, ask what the unit meant. The answer is itself information about how carefully the scope was read.
A related check: ask what happens to material that fails acceptance. If rejected hours are replaced at no cost, the per-hour figure already includes a rejection allowance, and comparing it against a quote without that clause is comparing two different products.
Make everyone quote the same scope
The cleanest method is to supply the quote template yourself. List the cost components as line items — recruitment, screening, consent documentation, equipment and studio, recording, annotation, quality control, delivery preparation — and require a number or an explicit exclusion for each.
The exclusions matter more than the inclusions. A quote that omits quality control is not cheaper than one that includes it. It is a quote for less work, and the difference is either paid later or discovered as a quality gap.
Fix the scope assumptions in the template as well: speaker count, hours per speaker, session structure, number of sites, languages, delivery date, and revision rounds. Quotes that assume different speaker counts are not variations on a price. They are prices for different projects.
Price the risk allocation, not just the work
Two quotes for the same scope can still differ because they allocate risk differently. Who bears the cost when a speaker drops out mid-project. Who pays if the first batch fails acceptance. Whether the price is fixed or re-quoted after the pilot. What a specification change costs and when that cost applies.
A quote with a lower unit price and an open-ended change process is not cheaper. It is cheaper in the scenario where nothing goes wrong, and few projects live there.
The practical approach is to convert the risk terms into a small set of scenarios — everything goes right, one rework round, one re-collection — and compare quotes across the scenarios rather than at the single optimistic point.
When unit price stops being the deciding number
There are conditions under which the per-unit figure carries almost no information.
- When the specification is unusual enough that only one producer can meet it. The negotiation is then about scope and schedule, not price.
- When the timeline is binding. A cheaper quote that lands after your launch date is not a cheaper quote.
- When the volume is small. At pilot scale the fixed costs dominate, and unit price differences are noise next to setup quality differences.
- When the risk terms differ. Price the scenarios, not the unit.
- When quality is the product. If the dataset is the differentiator, the comparison that matters is cost per accepted unit, and the rejected portion is paid for either way.
A scoring method that holds up
Once quotes are on a common basis, score them on more than price: fit to the specification, risk terms, schedule, quality process, and price — with the weights decided before the quotes are opened, not after.
Write down why each score was assigned. The written rationale is what makes the decision defensible to everyone who questions it later, whether that is a finance reviewer, a legal reviewer, or the team that inherits the dataset. It is also what makes the next procurement faster, because the reasoning stays reusable even when the producers do not.
The common failure here is scoring after the numbers are known, which reliably produces a ranking that matches the price list and calls it analysis.