Choosing a sample rate for a new collection

Record high, convert down, never the reverse. How to pick a rate for a new collection and verify what the recorder actually produces.

The sample rate is one of the few decisions in a data project that cannot be revisited. Audio recorded at 48 kHz converts down to anything below it with no loss that matters. Audio recorded at 16 kHz has no content above 8 kHz, and no amount of processing brings it back. The choice is between spending storage and spending a second recording budget.

The recommendation for a new collection is short: record at 48 kHz, 24-bit, one channel per microphone, and convert down later. The rest of this is the reasoning, the cases that change it, and how to check what your recorder is actually doing.

What 16 kHz gives up

Most recognition models consume 16 kHz audio, and 16 kHz is genuinely sufficient for the words themselves. What it gives up is the band above 8 kHz, where fricatives carry much of their energy and where speaker-characteristic detail lives. Two consequences follow. Fricatives sit right at the edge where the anti-aliasing filter is already rolling off, and any later work on speaker identity, anti-spoofing, or voice quality starts from a corpus that is missing the band those methods rely on.

For a corpus that will only ever feed a recognizer, 16 kHz is defensible. For a corpus meant to be reused, and collections usually outlive the project that funded them, 48 kHz keeps options open at a small storage cost.

What 48 kHz costs, stated plainly

Sixteen-bit mono at 48 kHz is three times the data rate of sixteen-bit mono at 16 kHz, and 24-bit makes it four and a half times. Across a large collection that is real storage and transfer cost. Set it against the rest of the project. Speaker recruitment, studio time, and annotation dominate the budget, and annotation is the line item you pay twice if the audio turns out to be insufficient.

If storage genuinely forces a trade, drop bit depth before sample rate. Sixteen bits at 48 kHz keeps the band and gives up headroom margin. Twenty-four bits at 16 kHz keeps the headroom and gives up the band permanently, and the band is the part that cannot be recovered.

Check what the recorder actually does

Device documentation is not a reliable statement of what happens to the signal. Phone applications and consumer recorders commonly report 48 kHz while resampling internally from 44.1 kHz, and a few high-quality modes record at 16 kHz and upsample on the way out, producing a file that claims a rate it does not contain.

Verify with a recording rather than a spec sheet. Record a broadband sound such as a hand clap, a key jingle, or a swept tone, and inspect the average spectrum. A hard cutoff near 8 kHz means the chain is really 16 kHz. A cutoff near 20 kHz means 44.1 kHz. Content above 22 kHz means the 48 kHz claim is honest. Run this once per device and app version, because the answer changes with both.

Rules that hold across a collection

Five rules, all cheap to follow at capture time and expensive to fix afterward.

  • One rate for the whole collection, and one rate within a session. Mixed rates inside a corpus are a reliable source of downstream bugs, and mixed rates inside one speaker session make that speaker difficult to compare with the others.
  • Never record at 8 kHz, even for a telephony deployment. Downsampling to 8 kHz later takes seconds; recording at 8 kHz cannot be undone.
  • If the device caps at 44.1 kHz, accept it and note it. The conversion from 44.1 kHz to 16 kHz is a rational resample that the standard tools handle correctly.
  • Keep the original files at the recorded rate. Convert copies for training and leave the archive untouched.
  • Write the rate, bit depth, and device into the session metadata at capture time, not into a spreadsheet that gets lost.

The decision, by case

Recognition only, telephony deployment: record at 48 kHz if the budget allows, otherwise 16 kHz, and never below. Recognition plus speaker or voice-quality work: 48 kHz and 24 bits. Reuse with an existing 8 kHz pipeline: still record high and convert, because a pipeline can accept a conversion and cannot accept the reverse. Tight storage: 16 kHz and 16 bits is the floor for a corpus used for recognition, and it should be treated as a constraint rather than a preference.

The cost of recording higher is paid once, in a currency that is easy to quantify. The cost of recording too low is paid later, and it is usually a second collection.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com