Questions to ask before buying a speech dataset, and the red flags in the answers

A dataset listing describes what it contains. These questions get at what it leaves out, and each group of answers has a pattern worth stopping for.

Catalog listings for speech datasets are written to sell, which means they describe hours, languages, and quality in the most favorable framing the data supports. The information that decides whether a dataset is usable for your project is mostly absent from the listing, and it does not appear until you ask.

The questions below are grouped by the risk they address. In each group, the answer to worry about is usually not a lie. It is a reasonable-sounding answer that happens to leave the risk unaddressed, and the most useful signal of all is how answers arrive: a seller who responds slowly, in writing, with specifics, including an occasional "that would need to be checked," is describing a dataset they actually know.

Provenance: where the audio came from

These questions establish whether the recording chain can be audited.

  • Who recorded it, in what country, and in what year?
  • What was each speaker told the recording would be used for?
  • Can the consent records themselves be reviewed, not just summarized?

Composition: who is actually in the audio

A dataset that cannot name a recording year was probably assembled from several sources with different consent standards, and "consent was obtained" without the text of the consent form does not tell you whether AI training was covered. A general media release usually is not. If consent documentation is treated as proprietary, the chain cannot be audited, which is the one thing you are buying.

On composition, a seller who will not produce a speaker manifest at the question stage will not produce one at the audit stage. A dataset where a few speakers dominate the total hours will overfit to those voices no matter how large the total is, and a demographic spread described as natural often means nobody controlled it.

  • How many distinct speakers are there, and how many hours does the largest single contributor account for?
  • What is the age, gender, and regional spread, and was it designed or did it happen?
  • How were speakers recruited? A call for volunteers selects for a different population than paid recruiting.

Conditions: how it was recorded

"Broadcast quality" is a category, not a measurement. If no signal-to-noise figure exists, the noise profile of the dataset is unknown, and noise robustness is one of the most common reasons a purchased dataset underperforms.

Variation across batches is the harder problem. A dataset recorded on three different setups has three different acoustic signatures, and a model trained on it may learn the setups instead of the speech.

  • What device, microphone, and sample rate were used?
  • Studio, office, or field — and if field, what was the measured noise floor?
  • Is the recording setup identical across all sessions, or does it vary by batch?

Annotation: how the text was produced

No agreement figure means annotation consistency was never measured, which is not the same as it being fine. "Native speakers" answers a different question than the one you asked, because nativeness does not imply training in a shared convention. And a guideline that is described but not shared usually contains conventions that will surprise you at the point where fixing them is expensive.

  • What does the annotation guideline say about filler words, numerals, and unintelligible speech?
  • Was any subset double-annotated, and what was the agreement rate?
  • Who annotated: trained annotators, native speakers, or crowdsourced workers?

Rights: what the license actually allows

A license that covers training but not distribution of the trained model will surface during product legal review, long after purchase. Sublicensing terms that are silent on derivatives usually mean the seller has not decided, and will decide against you when the question becomes real.

The withdrawal question is the sharpest test in this group. A seller with a considered answer has thought about their chain. A seller who has never considered it has a chain that can break, and the break will happen on a date you do not choose.

  • Is the license limited to training, or does it cover distributing a model trained on the data?
  • Can the data be used for benchmarking against competitors, or is that restricted?
  • If a speaker withdraws consent, who is responsible for removing their data downstream?

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com