How to choose an AI data vendor when there is no track record to check

A proposal tells you what a vendor can promise, not what they can deliver. Here are the dimensions that predict the difference, and how to test them early.

The training data market has no reliable reputation layer. There is no independent audit that tells you whether a vendor claiming forty languages can deliver in forty languages, and no public record of how their last project went. Every proposal arrives looking competent, because proposals are written by people whose job is to make them look competent.

Selection therefore has to be a process of finding out what a vendor cannot do, rather than confirming what they say they can. The questions below are built for that.

Separate the supply chain into its actual stages

Most vendors present themselves as one capability. In practice a data project runs through four stages, and a vendor may own all of them or subcontract some.

  • Recruitment: finding, screening, and scheduling the speakers or collectors.
  • Production: the recording itself, or the collection work in the field.
  • Annotation: turning raw material into labeled data under a written guideline.
  • Quality control: an independent check that can reject work before delivery.

Compare on four dimensions, not one

Ask who performs each stage and whether that party is in-house or a subcontractor. A vendor that owns recruitment and subcontracts annotation is not a problem by itself. A vendor that cannot say who annotates the data is a problem, because nobody in the chain is then accountable for the guideline being followed.

Capability, capacity, coverage, and accountability are separate properties, and vendors are usually strong on one and weak on another.

  • Capability: whether this exact kind of project has been run before, in the same modality, conditions, and annotation depth.
  • Capacity: how many projects run at once, and what happens to yours when a larger one arrives.
  • Coverage: whether the languages, regions, and recording conditions you need are reachable at all.
  • Accountability: who signs off on quality, and what happens contractually when a delivery is rejected.

Tests that work without a track record

A vendor with strong capability and thin capacity delivers a good pilot and a late production batch. Ask how many projects run concurrently, and what the largest current one looks like.

References are weak evidence, because vendors choose them. Three substitutes work better.

  • Ask for a failure story. Anyone who has run real projects has one that went wrong. How it was noticed, how fast, and what it cost is more informative than any success reference.
  • Ask for the annotation guideline. It exists as a document or it does not. If it does not, the conventions are being decided by each annotator individually.
  • Give a hypothetical. Describe a genuine conflict in your own specification, such as an utterance that is partly unintelligible, and ask how they would instruct an annotator.

What the proposal itself reveals

How a vendor responds to your specification, before any contract exists, previews how they will behave during it.

A vendor who engages with specifics — pushing back on a sample rate, asking which deployment the model targets, flagging that your speaker target is unreachable in your timeline — is doing the work that prevents rejected deliveries. A vendor who agrees with everything has either not read the specification or intends to resolve the disagreements later, at your expense.

Response speed matters less than response substance. A fast generic reply is worse than a slow reply that engages with the text.

Reasons to end the conversation

Some answers are informative enough to stop a procurement, regardless of the price attached to them.

  • They cannot describe how a speaker or collector gets rejected during screening.
  • They will not name their subcontractors or say which stages they subcontract.
  • The timeline contains no recruitment phase, as if speakers appear on day one.
  • They decline a paid pilot, or offer a free one produced by the sales team.
  • Quality is described only in adjectives, with no measurable acceptance criterion attached.

More insights

  • Handling a data batch that failed acceptance

    A failed batch is a decision with three possible outcomes, and the expensive mistake is choosing the wrong one. How to measure the failure, and what to agree beforehand.

  • When to stop collecting data

    More data stops helping before it stops costing. The signals that the curve has flattened, and the stopping rule that makes the decision mechanical.

  • Avoiding scope creep in a data project

    In a corpus project every property is additive, so requests arrive continuously and each one looks small. The fix is a classification step before any pricing.

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com