Sourcing speakers of a rare language
The constraint is not finding people who speak it. It is finding people who meet every other requirement at the same time, and proving that they do.
The constraint is the intersection
Finding speakers of a rare language is rarely the hard part. Finding speakers who meet every requirement at once is: the right variety, a recording environment quiet enough to use, availability inside the project window, literacy in the orthography the annotation will use, and enough of them to reach a speaker target rather than a token sample.
Each additional constraint cuts the pool multiplicatively, and the pool was small to begin with. That is why recruitment lead time rather than recording capacity sets the schedule on these projects, and why the estimate has to come from a probe rather than from a formula.
Channels, and what each one costs
The channels below run roughly from the highest quality per contact to the lowest, which is not the same order as speed or reach. Most projects end up working three or four of them at once, and the mix changes as the pool opens.
- Community organizations and cultural associations. The highest quality per contact and the slowest to convert. Access usually runs through one person whose trust has to be earned, and the introduction is often worth paying a local intermediary for.
- Diaspora communities in cities you can reach. Dense and easy to schedule, with one caveat to screen for: the variety may have drifted, and second-generation speakers may use the language in a narrower set of domains than the corpus requires.
- University language departments. Reliable for the standard variety and for people who can also annotate, with two drawbacks. Term-time availability is lumpy, and academic speakers are unrepresentative of the general population in education and register.
- Religious institutions and community radio. The best route to older speakers who are not reachable online, and it needs a phone-first process rather than an application form.
- Whoever produced the previous corpus in that language. The fastest route where it exists, because the speakers are already known to be able to do the work. Check how heavily that pool has already been recorded, since a heavily recorded pool may already sit inside public models.
- Referrals from the first speakers recruited. After three or four people are in, this becomes the best channel. Pay for the referral, and ask each speaker for people outside their household and outside their occupation, so the pool does not collapse into a single social cluster.
- Online language communities. Cheap to reach and the most expensive to verify, because the screening cost per accepted speaker is highest and the typical failure is a speaker of a closely related language rather than of the target.
Verification, because self-report is not evidence
Native speaker is a self-description with no standard behind it, and where a screening call is worth money, it will be claimed by people who do not meet it.
The screen that works is a short unscripted recording rather than an interview. Ask for two or three minutes of spontaneous speech on an open prompt, in the target language, on whatever device is at hand. Then have a reviewer from the target variety score it against a short rubric: vowel inventory and tone, stress placement, whether the everyday word or a loanword appears, morphology, and the discourse markers that mark a fluent speaker.
- Reading a passage tests literacy and script, not nativeness. It is useful as a separate check, and it selects for educated speakers, which may or may not be what the corpus needs.
- Build a variety probe: a short list of lexical and phonological items that separate the target variety from its neighbours. The list has to come from a speaker of the variety, not from a reference grammar.
- Use a domain switch test to find out what the speaker can do beyond the home. Ask them to describe something technical or professional in the language. Heavy code-switching at that point is a fact to record rather than a disqualification, but it changes which corpus the speaker fits.
- Put two independent reviewers on the first batch of screens and measure their agreement. If two reviewers from the variety cannot agree on who qualifies, the criterion is not defined well enough to apply consistently, and the fix belongs in the rubric.
- Keep every screen recording. It is evidence, it is the start of a speaker manifest, and it doubles as a sample of what the microphone will hear.
Decide about bilingual speakers first
For many rare languages the realistic pool is bilingual, and a purely monolingual speaker may be rare or nonexistent. That is not a defect in the pool, but it is a decision the specification has to make, because the answer changes the data.
Record the second language as a field for every speaker. If code-switching is acceptable or desirable, the guideline needs a rule for it, and that rule has to exist before the annotators meet the material. If it is not acceptable, the screen has to test for it, and the acceptable pool shrinks accordingly.
Discovering this at delivery is the expensive version, because the corpus has already been produced and no amount of annotation can repair the mismatch.
Literacy is a separate, smaller pool
Annotation work needs people who can write the language in a defined orthography, and that requirement is independent of whether they speak it natively. A fluent speaker who never writes the language is a good recording participant and a poor annotator.
Test it rather than assume it. Give candidates a short recording and ask for a transcription in the intended orthography, then score the result against the guideline: word boundaries, diacritics, how numerals are handled, and how an unintelligible stretch is marked. That score is the qualification, and the same test produces the first calibration data for the language.
Lead time, and keeping the pool warm
Recruitment lead time scales with distance from your existing channels, not with the number of speakers needed. Going from ten speakers to twenty in a pool you have already opened costs far less time than finding the first ten, which is why the probe matters more than the target.
Expect to screen several people for every one you accept, and build the schedule from that ratio rather than from the target count. The ratio is what converts a speaker target into calendar weeks.
- Pay properly for screening and for referrals. The expensive failure in a rare-language project is a scheduled session that produces unusable audio, and screening is what prevents it.
- Budget for no-shows and replacements explicitly. Confirmed attendance is weaker evidence in a pool reached by phone than in a studio booking system.
- Plan for a portable setup. Sessions in a rare language often happen in a home or a community space, so the equipment list and the quiet-room checklist have to travel.
- Keep the pool. The list of verified speakers, with consent for future contact recorded separately from consent for the recording itself, is the asset that makes the second project in that language cheaper than the first.