Transcription guidelines: writing down the rules for numbers, fillers and laughter

The guideline is the dataset. These are the decisions to settle before annotators start, and the calibration round that finds the ones you missed.

Verbatim or clean, decided once

The first fork in any transcription guideline is whether the transcript is verbatim or clean. A verbatim transcript keeps disfluencies: the fillers, the false starts, the repeated words. A clean transcript removes them and reads like written text. Either choice can be correct — recognition training usually wants the verbatim form, while uses such as synthesis or translation want the clean one — but mixing them inside a single dataset is not defensible, because the same behavior becomes an insertion in one file and correct text in the next, and no model can learn from that.

The choice belongs in the guideline, in the dataset documentation, and in the scorer. A buyer who scores a verbatim model against clean references will see insertions that are really conventions, and the reverse pairing hides real omissions. Writing the convention down is what makes the resulting number mean something.

Numbers and dates: write what was said

The rule that prevents most disagreements is short: transcribe the spoken form. Twenty twenty-six and two thousand and twenty-six are both said out loud, and the transcript records which one was spoken. Digits are a later, deterministic normalization step that a script applies to both the reference and the hypothesis at scoring time.

Putting digits in the transcript instead creates two problems. It breaks the match with the audio, which the aligner depends on, and it forces every annotator to invent a formatting convention, so the dataset ends up with 1,000, 1000 and a thousand as three spellings of one utterance. The same logic covers money, measurements and clock times: three thirty p m is what was said, and the conversion to a formatted time belongs to the normalizer rather than to the annotator.

Fillers, false starts and repeated words

If fillers are kept, their spellings must be fixed. Uh, um, er, hmm and mm each get one canonical spelling, or the vocabulary collects uh, uhh and uhm as separate tokens and the model spends capacity on spelling variation that carries no meaning.

  • Whether lengthened sounds are marked, and with what, such as a trailing colon on the stretched vowel.
  • How a false start is written: the abandoned fragment is transcribed as spoken, with a written convention for words cut off mid-sound.
  • Whether repetitions are kept. A repeated word such as "the the" is transcribed as said under a verbatim policy, and dropped entirely under a clean one.
  • What counts as a filler. A word such as "like" acts as a filler only in some positions, so define it by function and position, or keep every hesitation sound without judgment.

Laughter, coughs and everything else that is not a word

Non-speech events get a fixed tag set, applied the same way every time: laughter, cough, sigh, breath, noise, music, and the two that are most often confused, unintelligible and inaudible. The difference is worth a sentence in the guideline: unintelligible means speech is present but cannot be made out, while inaudible means there is nothing to hear. Without that distinction the two tags are used interchangeably and neither one is measurable.

Events that overlap words are tagged at their position and the words are kept, so a laughter tag sits beside the words that were spoken through it. Overlapping speech from a second person is tagged as overlap, and the guideline says whether only the primary speaker is transcribed or both are, because that choice turns a recognition set into something closer to a diarization set. Breath sounds are usually not transcribed, and saying so explicitly prevents thousands of small decisions. Words cut off at the edges of a clip get a marker rather than a guess.

Test the guideline before the whole batch meets it

A guideline is a hypothesis about what is clear, and the way to test it is a calibration round. Two or more annotators transcribe the same small set — ten to twenty clips chosen to hit the known trouble spots: numbers, fillers, laughter, overlapping speech, a name, a word from another language. Their outputs are compared with the same normalizer that will score the dataset, and every disagreement is either a missing rule or an ambiguous one. The fix goes into the guideline, not into the annotators.

Two artifacts come out of that round. The revised guideline, versioned, with the version stamped onto every delivered file so a later batch can be checked against the same rules. And the disagreement rate from the calibration, which is an honest floor for the dataset: if two careful annotators disagree on a small share of words, that share is the noise level inside the labels, and it belongs in the delivery notes rather than being discovered later by whoever trains on the data.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com