Accent labels that annotators can apply the same way twice
Accent is a continuum, so the label scheme is a lossy choice. Provenance beats perception for the category, and listeners are reliable only on strength.
The scheme is a lossy choice, so make it deliberately
Accent varies continuously, and any label set discretizes it. Two failure modes follow. A scheme that is too fine produces labels no two annotators apply the same way, and the dataset ends up with noise where a feature should be. A scheme that is too coarse — a single non-native bucket, for example — carries no information, and it cannot stratify a collection or support a result reported per accent.
The way out is to decide what the label is for. If it exists to balance a collection and to slice an evaluation, a small set of stable categories is enough. If it is meant as a training target for an accent classifier, the categories have to be learnable from audio alone, which rules out anything defined by geography rather than by sound.
Separate provenance from perception
There are three sources of truth about a speaker and they disagree. Where the speaker grew up and learned the language is a fact. How the speaker describes their own accent is a self-report. What a listener hears is a perception. The category label should be built from the first of those, because it is the reproducible one: for a second-language accent, the first language plus where the second language was learned and used; for a native variety, the region where the speaker grew up.
Provenance comes from the speaker, through a short intake form, not from an annotator listening to a clip. That form is part of the collection rather than an afterthought, and its fields are the ones that survive scrutiny: first language, region of upbringing, years in the current region, and the languages spoken at home. The perceived accent is then a derived value used for strength, never as the category itself.
What a listener can and cannot judge
Asking an annotator to name the variety of a speaker from a region they have never heard produces noise dressed as data. Asking for strength is a different task and a far more reproducible one: how far the pronunciation departs from the reference variety of the target language, on a short scale such as native, light, moderate, heavy. That judgment can be anchored with example clips and calibrated before production.
The calibration is the part usually skipped and the part that makes the labels usable. Give every annotator the same small anchor set, run a round, and measure how often they agree with each other on it. If agreement is low, the scale is wrong — coarsen it, or sharpen the definitions — rather than the annotators. Re-run the calibration after any change to the guideline, because the anchors define the scale.
Granularity rules that survive contact with data
Accent is a property of the speaker, so the label belongs on the speaker record rather than on each utterance. Per-utterance accent labels drift within a recording as the speaker shifts register or reads a different kind of sentence, and that drift is style, not accent.
- Region at the country level for native varieties. Anything finer needs speaker counts that a typical collection cannot support.
- First language for second-language accents, plus the strength level, rather than a guessed region of origin.
- Three or four strength levels, defined by example, with the anchor clips kept alongside the dataset.
- An explicit unknown or mixed bucket with written rules for when it applies, and the share of the dataset that lands in it reported. A large unknown share means the scheme is failing, and forcing a choice instead converts that failure into wrong labels.
- The label applies to the language being spoken. A bilingual speaker has one accent in each language, not one accent overall.
What the evaluation slice needs from the labels
A per-accent evaluation slice needs multiple speakers per accent. A slice with one speaker per accent measures the speaker, and it will report differences between people as differences between accents. Report the speaker count beside every slice result, and keep the same speakers out of the training set.
Finally, record where each label came from: provenance, derived, or listener-judged. A buyer who knows which labels are facts and which are perceptions can use them differently, and a redelivery that adds speakers can be merged without guessing which scheme was in force. The label provenance field is small, and it is the difference between a dataset that can be extended and one that has to be rebuilt.