Auditing a dataset before you buy it: a sample-based checklist that finds the real defects
A listing describes the best of a dataset; an audit looks at the actual files. What to request, what to compute, what to compare, and how to end with a decision.
Get the right sample first
The audit is only as good as the sample it runs on, and the sample is a negotiating point. Ask for it in writing, with the selection method stated: a random draw across the whole dataset, not a curated folder. A few hundred items, or a low single-digit percentage of the corpus, is enough to find systemic problems. Curated samples find nothing, because curation is the thing being audited.
Ask for the sample to arrive with the artifacts that travel with the data: the manifest or metadata table, the annotation guideline in force, the consent or license documentation, and any quality report the seller keeps. A seller who cannot produce a guideline is usually not hiding one; there is no guideline, and that fact alone changes the rest of the audit.
Listen to the sample against a fixed checklist
Listen to every item in the sample at least in part, and score each against the same list rather than against an impression. Impressions are not comparable across sessions or across reviewers.
- Signal: noise floor in the quiet stretches, clipping on the loud ones, hum or electrical interference, and whether levels are consistent from file to file.
- Speech: crosstalk or background voices, whether the audible conditions match the claimed ones — a quiet office that contains a keyboard and a street — and whether the speaker matches the description in the manifest.
- Recording continuity: abrupt cuts mid-word, files that begin after the speech has started, duplicated content across files, and pieces of one session sold as separate items.
- Consistency: does file ten sound like it came from the same room and equipment as file one? A dataset assembled from several batches reveals its seams in the audio.
Check the metadata against the files
The manifest is where a dataset tells the truth about itself, because it is harder to curate than the audio. Take twenty rows at random and verify each against its file: duration, speaker id, transcript, tags.
- Counts: does the manifest hold the number of items the listing claims, and does the sum of durations match the headline hours — including whether "hours" means audio hours or wall-clock session time.
- Speaker concentration: the distribution of hours per speaker, and the largest contributor's share. This is often the single most revealing number in the whole audit.
- Missing values: which fields are incomplete, and whether a blank means unknown, declined, or not applicable. A column that is half empty is a column nobody can filter on.
- Text quality, where transcripts exist: read a page of them. Uniformly perfect punctuation across spontaneous speech suggests machine transcription presented as human work; human transcripts have a characteristic inconsistency. Either can be acceptable if the listing says which one it is.
Check consistency across time
A dataset built over months has a history, and the history shows up in the data. Group the manifest by month or batch, and look at what moves.
- Label or tag distributions per batch: a tag that appears in month one and vanishes, or a category that triples in month six, is a guideline change, a different annotator team, or a different source. The buyer should know which.
- Recording signatures per batch: device, sample rate, room tone. Batches recorded on different setups are different datasets wearing one cover.
- Annotator or team fields, where they exist: do the conventions hold across them, or does each team mark fillers its own way?
- Version fields: the guideline version recorded per item, where available, tells you how many annotation regimes the corpus contains.
Check the rights chain
The technical audit is half the job; the other half is paper. Request the actual documents rather than a summary: the consent form text as it was signed, the license terms offered to you, and the rights basis per source where the corpus is assembled from several. Two failure modes are common enough to check specifically — a consent form whose text never mentions AI training, and a chain of custody that breaks at an acquisition nobody can document.
Write the audit note, and end with a decision
Keep the output short enough that the person making the purchase decision will actually read it. The note ends with a recommendation in one of three forms — proceed, proceed with a defined rework or adjustment, or stop — and that recommendation matters most: an audit that ends in a pile of observations without a decision is a cost with no benefit.
Four parts carry the rest of the document:
- Per section: what was checked, on how many items, and what was found — pass, pass with exceptions, or fail.
- Exceptions with severity. A small rate of corrupt files is a commercial question; a consent form that does not cover training is a stop.
- The unresolved questions: what could not be verified from the sample, written as questions to the seller with the answer needed in writing.
- The findings that drove the recommendation — the two or three items anyone re-reads six months later, when the dataset is in training and something unexpected turns up.