Datasets by language and category

468 combinations, each with the collection problem that is specific to it. Nothing here is in stock — every entry describes what we source to order, and what your specification needs to settle before production starts.

Why these are pages and not products

A catalog listing answers "what do you have?" Our model answers "what do you need?" So there is no cart here and no price. What each page does carry is the part that does not change between projects: the specific thing that goes wrong when you collect this kind of data in this language, and the field in your specification that decides the quote.

That is deliberate. If you are comparing suppliers for, say, Thai call center audio, the useful question is not who has a file on a shelf — it is who knows that Thai script carries no word spacing and that your transcript convention therefore has to be agreed before recording, not after.

Call Center Speech

26 languages

  • Hindi Call Center Speech

    The dialect background of every agent and every caller has to be recorded person by person (Delhi, Uttar Pradesh, Bihar, and Rajasthan accents differ noticeably), otherwise one batch ends up mixing incompatible sound systems.

  • Arabic (MSA) Call Center Speech

    Every agent's native dialect must be annotated (Egyptian / Gulf / Levantine / Maghrebi), along with the range of accent deviation that is acceptable.

  • Indonesian Call Center Speech

    Whether spoken numbers and amounts are normalized to digits — Indonesian agents compress digit strings ('tiga lima' for 35), and the convention cannot be recovered from the audio afterwards.

  • Thai Call Center Speech

    The word boundary convention (as written / as spoken / delivered after word segmentation) belongs in the contract annex, and the three kinds of output must not be mixed in one batch.

  • Turkish Call Center Speech

    Agent origin has to separate Istanbul accent from eastern accent, and the two cannot be collected together, because model generalization differs sharply between them.

  • Vietnamese Call Center Speech

    Every recording gets the speaker's dialect region annotated (northern / central / southern), and delivery is split into batches by dialect region.

  • Filipino Call Center Speech

    Be explicit about whether the delivery is "pure Filipino" or "raw Taglish as spoken" — these are two different products, and choosing the wrong one wastes the whole batch.

  • Persian Call Center Speech

    Whether agent and caller are delivered on separate channels — Persian drops subject pronouns freely, so without a channel split the transcript cannot attribute an utterance to a speaker.

  • Tamil Call Center Speech

    Speaker region must be annotated record by record (Tamil Nadu / Sri Lanka / Malaysia / Singapore), with the share of each region stated.

  • Telugu Call Center Speech

    State which dialect region is the baseline, and allow Telangana dialect words into the word list; if they are not allowed, the annotation guideline has to carry an explicit whitelist.

  • Malayalam Call Center Speech

    Both speakers and annotators must be deduplicated, and delivery must include an annotator ID mapping (no identity disclosed, but batch membership traceable).

  • Nepali Call Center Speech

    Speaker metadata must record nationality, native dialect region, and language of education for every record — those three decide whether the batch can support accent stratification experiments at all.

  • Sinhala Call Center Speech

    Decide whether English loanwords are kept, tagged as switches, or replaced with Sinhala equivalents — the three produce completely different data.

  • Bangla Call Center Speech

    Specify which market is the baseline (Dhaka / Kolkata), and deliver in separate regional batches rather than one mixed batch.

  • Khmer Call Center Speech

    Whether the honorific register is transcribed as spoken or normalized — Khmer uses separate honorific vocabulary for the same verb, and agents shift register mid-call.

  • Arabic (Egyptian) Call Center Speech

    State whether the transcript is produced from the audio or derived from the script, and keep the two speakers on separate channels — the scripted standard-language half and the caller's Egyptian half have to stay splittable, and they cannot be separated once mixed.

  • Arabic (Gulf) Call Center Speech

    Annotate every agent as either a native Gulf speaker or a second-language Arabic speaker, with their first language recorded; deliver agent and caller channels separately so the two sound systems are never averaged together.

  • Kannada Call Center Speech

    Annotate the caller's dialect region on every call, and state whether the caller channel keeps colloquial forms or is normalized to the written standard — normalizing it erases the register split this data exists to capture.

  • Punjabi Call Center Speech

    Every agent and caller gets their side of the border and their dialect annotated, and the transcript script is fixed in the contract annex — a batch transcribed in the wrong script cannot be converted afterwards without a full re-annotation.

  • Burmese Call Center Speech

    Name the authority dictionary that fixes spelling in the contract annex, and state whether the transcript keeps the text as written or is normalized into word-segmented form — the two outputs cannot be mixed in one batch.

  • Amharic Call Center Speech

    The honorific convention belongs in the contract annex: whether polite verb forms are transcribed as spoken or normalized to the plain form. Amharic marks respect on the verb itself, so normalizing it erases the social signal the agent was managing.

  • Ukrainian Call Center Speech

    The contract annex has to state whether pure Ukrainian only is accepted or surzhyk is allowed, and if it is allowed, every recording carries a mixing-degree label.

  • Swahili Call Center Speech

    Every agent and caller needs two fields recorded person by person: first language (Kikuyu, Luo, Luhya, Sukuma, Chaga) and the market the call belongs to (Nairobi / Dar es Salaam). Without both, a batch labeled simply "Swahili" cannot be stratified at all.

  • Kurdish Call Center Speech

    The branch (Kurmanji or Sorani) and the transcript script are fixed in the order, and every agent and caller carries a country-of-origin tag, because the loanword layer tracks the state language on each side.

  • Hausa Call Center Speech

    Every speaker's first language and Hausa variety (Kano, Sokoto, Katsina, or a Nigerien variety) has to be annotated per speaker rather than per call; first- and second-language Hausa in one batch cannot be separated afterwards.

  • Uzbek Call Center Speech

    The transcript script convention (Latin or Cyrillic) and the rule for Russian-language segments both belong in the contract annex, and the two cannot be mixed inside one batch.

About Call Center Speech →

Conversational Speech

26 languages

  • Hindi Conversational Speech

    Whether overlapping speech is kept or split, and whether transcription uses Devanagari or Perso-Arabic script — both get settled when the order is placed.

  • Arabic (MSA) Conversational Speech

    Every speaker's dialect background and permitted deviation must be annotated individually, with the role cards attached verbatim, or the data cannot be reproduced.

  • Indonesian Conversational Speech

    The share of each language component in every conversation has to be annotated as a graded band, and delivery is stratified by that share.

  • Thai Conversational Speech

    Whether particles are kept or tagged as ignorable markers is decided by the downstream use, and it gets written into the delivery notes.

  • Turkish Conversational Speech

    Whether backchannel words (evet, hı hı, tabii) are transcribed as words or tagged separately — they are dense in Turkish conversation and handled inconsistently across annotators.

  • Vietnamese Conversational Speech

    The pronoun transcription convention must be explicit (as actually spoken / mapped to the standard equivalent), with a relationship reference table attached.

  • Filipino Conversational Speech

    Whether switch points are annotated at word level or sentence level, and whether English words are tagged separately — those two decide the order of magnitude of the work.

  • Persian Conversational Speech

    The transcription standard must state "transcribe the actual spoken pronunciation", and speakers' language of education must be annotated.

  • Tamil Conversational Speech

    Whether transcription follows spoken form or written form has to be explicit; spoken form is recommended, with a spoken-to-written comparison table for common words.

  • Telugu Conversational Speech

    Both speakers' dialect regions get annotated record by record, and cross-dialect conversations are delivered in separate groups from same-dialect ones.

  • Malayalam Conversational Speech

    The segmentation convention (by pause / by fixed duration / by semantic unit) gets fixed in advance, and output produced under different conventions cannot be mixed in one batch.

  • Nepali Conversational Speech

    Honorific forms are transcribed as spoken and never normalized, and the social relationship between the two speakers is annotated in the metadata.

  • Sinhala Conversational Speech

    The handling of English loanwords (kept / tagged / replaced) has to be locked, along with whether loanword share is reported separately.

  • Bangla Conversational Speech

    Both speakers must come from the same regional standard, noted in the metadata; cross-region conversations are grouped separately.

  • Khmer Conversational Speech

    Whether register is annotated at all, and which convention governs word boundaries — both have to be written into the annotation guideline.

  • Arabic (Egyptian) Conversational Speech

    The guideline must fix transcription as write-as-spoken, never corrected toward Modern Standard Arabic, and carry a worked example list of the highest-frequency pairs ('izzayy' / 'kayfa', 'eeh' / 'maadha', 'fiin' / 'ayna') so annotators have something concrete to check against.

  • Arabic (Gulf) Conversational Speech

    Annotate each speaker's variety separately (Saudi Eastern Province / Najdi / Kuwaiti / Bahraini / Emirati / expatriate second-language Arabic), and state whether mixed-variety pairs are accepted or must be held as separate groups.

  • Kannada Conversational Speech

    Provide a canonical colloquial spelling table covering reduced final vowels and contracted verb endings, and treat it as part of the guideline — without it, inter-annotator agreement on Kannada conversation cannot be measured.

  • Punjabi Conversational Speech

    Fix whether tone-bearing words are written with a tone mark or a phonetic note, and annotate both speakers' dialect region plus side of origin — a transcript without the tone column cannot be aligned back to the audio reliably.

  • Burmese Conversational Speech

    Fix the convention for reduced syllables (write the full form / write what was said / mark as reduced) and state whether overlapping turns are kept or split — neither decision can be revised once annotation is under way.

  • Amharic Conversational Speech

    Whether the transcript follows the written form or the spoken form, and if written, how dropped sixth-order vowels are marked. The two conventions produce texts that cannot be merged after the fact.

  • Ukrainian Conversational Speech

    Each speaker gets two fields: self-declared everyday language, and the observed language composition of the recording, measured separately. The two are never merged into one.

  • Swahili Conversational Speech

    Whether transcription follows standard orthography or records the spoken colloquial form, and whether English and local-language insertions are tagged — the two decisions have to be made together, or half the batch comes back normalized.

  • Kurdish Conversational Speech

    The branch is stated per recording, and every conversation carries the speaker's country and the state language mixed into it; Kurmanji and Sorani material is delivered as separate batches and never pooled.

  • Hausa Conversational Speech

    Whether tone and vowel length are annotated at all, and if so under which convention (a separate tone layer, or diacritics on the Boko text); the two conventions are not interchangeable within one batch.

  • Uzbek Conversational Speech

    Annotate every speaker's age band and script of schooling, and settle whether Russian insertions are transcribed in Cyrillic or transliterated into Latin — that choice cannot be recovered from the audio afterwards.

About Conversational Speech →

Singing Voice

26 languages

  • Hindi Singing Voice

    Whether vocal separation is applied (isolated vocal) or the mix is kept, and whether pitch / MIDI annotation is attached — those two decide the use and the cost.

  • Arabic (MSA) Singing Voice

    The genre scope (classical / anthem / modern) and whether dialect words may be mixed in have to be locked when the project is scoped.

  • Indonesian Singing Voice

    Whether English passages in the lyrics get their own language tag, and how phoneme boundaries are handled at the points where the language changes.

  • Thai Singing Voice

    Two annotations are needed for every syllable, the lexical tone of the lyric and the actual sung pitch, along with a statement that the two not agreeing is normal.

  • Turkish Singing Voice

    Whether alignment granularity is syllabic or phonemic, and how a long word spanning several notes is marked, need worked examples in the annotation guideline.

  • Vietnamese Singing Voice

    Each singer's regional accent gets annotated record by record, and delivery is batched by accent.

  • Filipino Singing Voice

    Tag the lyrics with a language label by line, and state the alignment rule for the points where the language changes.

  • Persian Singing Voice

    Separate classical poetry lyrics from modern colloquial lyrics into two classes, each using a different phoneme annotation table.

  • Tamil Singing Voice

    Source material must be original or already cleared, with the rights chain documented; the phoneme annotation has to state how weakened consonants are handled.

  • Telugu Singing Voice

    The register of the lyrics (colloquial / Sanskritized / mixed) has to be annotated, and delivery is stratified by register.

  • Malayalam Singing Voice

    Alignment effort has to be estimated in advance, and delivery is syllable-level alignment; whole-utterance timestamps alone are not accepted.

  • Nepali Singing Voice

    Folk songs and modern songs go into separate groups, each with its own statement of phoneme coverage.

  • Sinhala Singing Voice

    Mixed-language sections get their own language annotation and their own phoneme table; do not force them into the Sinhala phoneme set.

  • Bangla Singing Voice

    Literary repertoire and colloquial repertoire go into separate groups, each with its own phoneme table.

  • Khmer Singing Voice

    Traditional and pop go into separate groups; for the traditional group, state additionally whether ornaments are annotated separately.

  • Arabic (Egyptian) Singing Voice

    Whether the lyric sheet used for alignment follows what is actually sung or the published orthography — Egyptian songs are sung in dialect but printed in normalized spelling, and the two disagree on exactly the high-frequency words that alignment depends on.

  • Arabic (Gulf) Singing Voice

    Fix the lyric orthography convention for Gulf dialect words (dialect spelling versus normalized standard forms) and whether delivery is a dry vocal or the full mix with accompaniment.

  • Kannada Singing Voice

    Classify each item by repertoire register (devotional / poetry-based light song / film) and state which archaic forms the phoneme table covers — annotating archaic lyrics with a modern colloquial table produces systematic errors.

  • Punjabi Singing Voice

    State which script the lyric annotation uses, how English and Urdu lines are tagged, and whether each syllable carries the lexical tone of the lyric alongside the sung pitch — the two not agreeing is normal and has to be documented.

  • Burmese Singing Voice

    Label every syllable twice — the lexical tone of the lyric and the actual vocal production — and state in the delivery notes that stylistic creak or breathiness is not the lexical creaky tone.

  • Amharic Singing Voice

    Separate the traditional repertoire from modern pop, and state whether the vocal ships isolated or with the accompaniment. Lyric alignment has to state how double-meaning lines are transcribed and who arbitrates them.

  • Ukrainian Singing Voice

    For every song, state whether the lyrics are annotated for dialect and archaic vocabulary, and mark where the melody has moved the lexical stress away from its spoken position.

  • Swahili Singing Voice

    Whether the delivery separates the vocal from the accompaniment, and whether each syllable is annotated with both the lyric text and its actual sung duration — the melody overrides the length contrast the lyrics themselves rely on.

  • Kurdish Singing Voice

    State the branch and the regional singing tradition the material comes from, and deliver lyric transcripts in a fixed script, with a written note on how borrowed words in the lyrics were transcribed.

  • Hausa Singing Voice

    Two annotations per syllable, the lexical tone of the lyric and the actual sung pitch, plus an explicit note that the two disagreeing is normal; and whether Ajami source lyrics are transliterated into Boko for alignment.

  • Uzbek Singing Voice

    State the lyric script used for annotation and the alignment convention for melismatic passages — one syllable stretched across a melodic phrase will otherwise be segmented differently by every annotator.

About Singing Voice →

Speech Commands

26 languages

  • Hindi Speech Commands

    The word list has to cover the Hindi / English equivalent forms, and state the minimum speaker count and repetitions per speaker for every entry.

  • Arabic (MSA) Speech Commands

    Speakers have to cover the main dialect regions (Egyptian / Gulf / Levantine / Maghrebi), and delivery is stratified by dialect region.

  • Indonesian Speech Commands

    The pronunciation convention for loanword entries (Indonesianized / English original / both) has to be locked, and no words may be added after the lock.

  • Thai Speech Commands

    Every entry has to state whether it includes a politeness variant, and whether variants count toward the word list size.

  • Turkish Speech Commands

    The inclusion boundary (roots only / common variants / everything) has to be locked, with a table of variant generation rules provided.

  • Vietnamese Speech Commands

    The word list has to state which speaker age bands it covers, and each age band needs enough speaker samples.

  • Filipino Speech Commands

    Entries are defined separately for each mixing pattern (pure Filipino / pure English / mixed), with the share of each stated.

  • Persian Speech Commands

    The word list uses spoken form, with the written equivalents attached for reference, and recording does not follow written-form pronunciation.

  • Tamil Speech Commands

    The word list has to name its source region, and state whether cross-region variants are covered.

  • Telugu Speech Commands

    The word list has to state its dialect baseline, and speaker recruitment has to cover the main dialect regions.

  • Malayalam Speech Commands

    The minimum speaker count and repetitions per speaker go into the contract for every entry, and padding with a few speakers doing many repetitions is not accepted.

  • Nepali Speech Commands

    Whether the word list ships with a recording script and a phonetic spelling for every entry — Nepali has no single romanization convention, so each collector invents their own.

  • Sinhala Speech Commands

    English loanword entries get their own section, with the urban and rural speaker distribution stated.

  • Bangla Speech Commands

    The word list is defined and delivered per region, and the two are not merged into one Bangla word list.

  • Khmer Speech Commands

    The word list carries two columns, the orthographic spelling and the phonemic transcription, and recording follows the phonemic column.

  • Arabic (Egyptian) Speech Commands

    The word list carries an orthographic column and a phonemic column side by side, and states whether nativized English commands (play / stop / cancel) count as entries — they are high-frequency in Egyptian use and cannot be added once the list is locked.

  • Arabic (Gulf) Speech Commands

    Ship a pronunciation lexicon with the word list, marking every entry whose Gulf pronunciation departs from the written form, and record each speaker's region so that /ɡ/, /tʃ/, and ج variation stays traceable.

  • Kannada Speech Commands

    Every entry declares its imperative register (familiar / polite), and the list states how English-plus-light-verb hybrids are counted — one entry or two. Both decisions are made before recording starts.

  • Punjabi Speech Commands

    Declare the baseline side on the word list itself, carry a tone column for every entry (tone is unwritten in both scripts), and fix the speaker count and repetition targets before recording starts.

  • Burmese Speech Commands

    Deliver the word list in a single declared text encoding, proofread after conversion, with the colloquial imperative as the recorded form and the literary equivalent attached for reference only.

  • Amharic Speech Commands

    The word list has to state its ejective minimal-pair coverage and carry the polite and plain imperative variants as separate entries. The list is locked before recording and nothing is added after the lock.

  • Ukrainian Speech Commands

    The word list has to state whether the Russian-derived command forms users habitually say are covered as separate entries, and carry a stress-marking column for every entry.

  • Swahili Speech Commands

    Whether entries are written in standard imperative form or in the colloquial forms users actually speak, and whether English equivalents are separate entries — the list has to be locked before recording, and no entries may be added midway.

  • Kurdish Speech Commands

    The command list is frozen per branch and per country before recording, with its written form fixed in one script; a list carried over from the other branch cannot be recorded as it stands.

  • Hausa Speech Commands

    Every entry ships with tone and vowel-length marking (diacritics or a phonemic column), and the recording script cannot be plain orthography; the ambiguity is invisible on the page and will not surface until after recording.

  • Uzbek Speech Commands

    Every command entry needs both the Russian-derived and the native coinage variant, with the age distribution of speakers recorded per entry — the word list cannot be locked until both forms are covered.

About Speech Commands →

Read Speech

26 languages

  • Hindi Read Speech

    Whether recording happens in a studio (low noise floor) or at home, and whether an extra batch at natural speaking rate is collected alongside — those two decide the cost and the use.

  • Arabic (MSA) Read Speech

    The text must supply full vocalization (short vowels and case endings included), and the rule for choosing among words with multiple valid readings has to be stated.

  • Indonesian Read Speech

    Speakers' language of education and native dialect region have to be annotated record by record, so that pronunciation differences can be explained.

  • Thai Read Speech

    Whether word segmentation (word boundaries) is attached has to be explicit, because the effort is substantial and it affects the delivery timeline.

  • Turkish Read Speech

    Whether stress position is annotated; if it is, the stress system used has to be stated.

  • Vietnamese Read Speech

    The text must keep its diacritics intact; speaker dialect region is annotated record by record, and northern and southern speakers are never mixed in one batch.

  • Filipino Read Speech

    Whether an extra batch at natural speaking rate is collected as a comparison has to be decided in advance; the comparison batch ships separately and is not mixed in.

  • Persian Read Speech

    Whether the ezafe linker is transcribed — it is unwritten in Persian script, so the source text gives no ground truth and different annotators will mark it differently.

  • Tamil Read Speech

    Tamil orthography does not mark the voiced allophones of its plosives, so the same written text is read differently by different speakers — fix whether the transcript records the written letter or the pronounced sound.

  • Telugu Read Speech

    The text has to supply the phoneme sequence after glyph decomposition, otherwise it cannot be used directly for acoustic model training.

  • Malayalam Read Speech

    Speaker count and duration requirements have to be locked in advance, so standards are not quietly lowered mid-project because recruitment fell short.

  • Nepali Read Speech

    The source and rights status of the reading text have to be documented, and text of unknown origin is not accepted.

  • Sinhala Read Speech

    The reading convention for loanwords has to be explicit, and loanwords should be marked in the text itself.

  • Bangla Read Speech

    Text and speaker must be matched to the same region, and delivery is batched by region with the corresponding text version attached.

  • Khmer Read Speech

    Whether the reader must pronounce subscript (coeng) clusters in full or may reduce them as in casual speech — read-speech corpora fail on exactly this inconsistency.

  • Arabic (Egyptian) Read Speech

    Whether the source text is fully vocalized and whether readers are instructed to use Egyptian or standard vowel choices — the short vowels are not in the writing system, so the reader's own dialect decides them and the same text yields different audio from different readers.

  • Arabic (Gulf) Read Speech

    State whether the text is standard or Gulf dialect, and if dialect, attach the spelling convention as a contract annex; also state whether dialect coloring in a standard-language reading is accepted or must be re-recorded.

  • Kannada Read Speech

    Deliver the reading text normalized to one stated Unicode form, and run a normalization pass across every text source before recording; list any item whose text could not be normalized in the alignment report.

  • Punjabi Read Speech

    Lock the reading script and the speaker side as one matched pair, and state that tone annotation is derived from the audio alone — the text version cannot serve as tone ground truth.

  • Burmese Read Speech

    Attach a word-segmented, pronunciation-annotated version of the reading text so the audio can be aligned, and state whether the recording follows the literary reading register or ordinary colloquial pronunciation.

  • Amharic Read Speech

    Whether the reading text carries a gemination decision for ambiguous words, and whether the reader is allowed to choose. Left open, two readers deliver two different corpora under one text.

  • Ukrainian Read Speech

    Deliver a stress-marked version of the reading text as an attachment, and state which stress convention the readers were given wherever more than one variant is accepted.

  • Swahili Read Speech

    Speakers' first-language background has to be annotated record by record even in read speech, and the studio-versus-home choice has to be fixed when the order is placed — the two differ in noise floor and in cost.

  • Kurdish Read Speech

    The alphabet and spelling standard for the prompt text are fixed in writing and attached to the delivery, and readers are screened for literacy in that script before recording rather than after.

  • Hausa Read Speech

    Annotate the reader's variety per record and deliver in batches by variety; and decide whether the reading text ships with tone and vowel-length marking, since plain Boko leaves both entirely to the reader.

  • Uzbek Read Speech

    The reading text must be authored in one script and that choice stated (Latin is the practical default); text converted from Cyrillic has to be proofread by a native reader before recording.

About Read Speech →

Podcast Speech

26 languages

  • Hindi Podcast Speech

    Whether diarization annotation is needed — the most labor-intensive item in long-form audio — has to be decided first.

  • Arabic (MSA) Podcast Speech

    Annotate the language variety segment by segment (MSA / each dialect), and state the share of total duration for each variety.

  • Indonesian Podcast Speech

    Annotate language makeup by segment rather than by episode, and state the segment granularity.

  • Thai Podcast Speech

    Segment length (say 10 seconds or 30 seconds) and whether overlap is annotated — both have to be locked in advance.

  • Turkish Podcast Speech

    Whether the loanword rule is applied per host or per episode — Turkish podcast hosts differ enough in how they pronounce borrowings that one rule across a series produces inconsistent transcripts.

  • Vietnamese Podcast Speech

    Each host gets a separate dialect region annotation, and the pronoun transcription convention has to be explicit.

  • Filipino Podcast Speech

    Annotate language by segment, and state whether English stretches are cut into separate files.

  • Persian Podcast Speech

    The transcription standard has to state that spoken form is the default, and set out how quoted passages are handled.

  • Tamil Podcast Speech

    The register distribution has to be measured and delivered (how much spoken form versus written form), and delivery is stratified by register.

  • Telugu Podcast Speech

    Every speaker's dialect region is annotated record by record, and long-form audio is delivered in batches by dialect region.

  • Malayalam Podcast Speech

    Annotator IDs have to be supplied with the delivery (no identity disclosed), so that batch membership can be traced and deduplication verified.

  • Nepali Podcast Speech

    The source of material (licensed public audio / original recording) has to be documented, and honorific forms are transcribed as spoken rather than normalized.

  • Sinhala Podcast Speech

    A loanword share statistic has to be delivered, along with a statement of whether loanwords are tagged separately.

  • Bangla Podcast Speech

    Whether episodes are delivered segmented and whether a per-episode speaker roster is included — Bangla podcasts are mostly multi-host, and without the roster the speaker labels are unusable.

  • Khmer Podcast Speech

    Whether register annotation is in scope has to be decided in advance; if it is, a register comparison word list has to be provided.

  • Arabic (Egyptian) Podcast Speech

    Annotate the speaker's region within Egypt for every host and guest, and mark the passages read aloud from written material as standard-language segments — they alternate with colloquial talk inside a single episode and one language label hides both.

  • Arabic (Gulf) Podcast Speech

    Annotate variety per speaker — host and each guest separately, by Gulf country — and fix the convention for English loanwords (transcribed as pronounced, tagged, or replaced with an Arabic equivalent).

  • Kannada Podcast Speech

    Decide whether scripted segments (sponsor reads, quoted material) are cut and labelled separately from improvised talk — they are read rather than spoken, and mixing them into the episode statistics hides the register split.

  • Punjabi Podcast Speech

    Annotate the production origin of each episode and fix the transcript script for the whole series — pooling episodes from different production centers under one language label hides the mix difference that drives everything downstream.

  • Burmese Podcast Speech

    Measure and deliver the register distribution of every episode (colloquial versus literary share), and confirm the team location and data handling arrangement in writing before annotation begins.

  • Amharic Podcast Speech

    Name a single spelling authority in the guideline, one dictionary or one published norm, and state whether the same annotator covered a whole episode. Spelling habits drift between people and long-form audio shows it.

  • Ukrainian Podcast Speech

    Language is annotated per segment, not per episode, with a mixing-degree band attached to each segment; whether diarization is required gets settled at the same time.

  • Swahili Podcast Speech

    Language and register are annotated per segment rather than per episode, with the segment length stated; long-form audio also needs a speaker roster, since most Swahili podcasts run two or more hosts whose turns cannot be separated from the transcript alone.

  • Kurdish Podcast Speech

    Each episode carries the speaker's branch, country of residence, and the second language mixed in; whether those segments are transcribed, tagged, or excluded is fixed before annotation begins.

  • Hausa Podcast Speech

    Decide whether tone and length annotation is in scope for long-form audio at all, and if it is, set a fixed QC sampling rate per recorded hour so that marking consistency can be measured instead of assumed.

  • Uzbek Podcast Speech

    Annotate each speaker's country of origin, generation, and script of schooling, segment by segment — long-form Uzbek audio routinely mixes speakers whose Uzbek is not the same Uzbek.

About Podcast Speech →

Multilingual Speech

26 languages

  • Hindi Multilingual Speech

    Speaker count, hours, and recording conditions have to be listed language by language. A single project-wide total will not do.

  • Arabic (MSA) Multilingual Speech

    The granularity of the Arabic labels — merged or split — has to be locked when the project starts. Changing it midway destroys annotation that has already been done.

  • Indonesian Multilingual Speech

    Provide decision rules and examples for telling Indonesian from Malay, and state whether local languages get their own labels.

  • Thai Multilingual Speech

    State whether Thai runs through a separate word segmentation step, and name the segmentation tool and its version number.

  • Turkish Multilingual Speech

    Loanwords must be annotated against the pronunciation rules of their own language. Do not share one cross-language phoneme inventory.

  • Vietnamese Multilingual Speech

    Vietnamese has to be split into sub-labels by dialect region, with the sample size for each region stated.

  • Filipino Multilingual Speech

    Set one project-wide rule for handling mixed speech, and state the English share separately for Filipino.

  • Persian Multilingual Speech

    State whether Persian, Dari, and Tajik get separate labels, and annotate each speaker's country of origin item by item.

  • Tamil Multilingual Speech

    How the language of each file is labelled when Tamil is collected alongside Malayalam — the two are close enough that automatic language ID mislabels them, and a wrong label is not recoverable downstream.

  • Telugu Multilingual Speech

    Define the geographic label system separately for each language. Do not force them into one scheme.

  • Malayalam Multilingual Speech

    Assess Malayalam's capacity and timeline separately. Do not assume it will deliver on the same schedule as the other languages.

  • Nepali Multilingual Speech

    Provide decision rules and boundary examples for Nepali versus Hindi, and spot-check language label accuracy during QC.

  • Sinhala Multilingual Speech

    State whether Sinhala-Tamil mixing gets its own label, and report the share of English content.

  • Bangla Multilingual Speech

    Which national orthography the transcription follows — Bangladesh and West Bengal spell the same word differently, and pooling the two in one corpus splits the model's training signal.

  • Khmer Multilingual Speech

    Assess Khmer's capacity separately and write it into the schedule. Do not estimate it from the project average.

  • Arabic (Egyptian) Multilingual Speech

    Report the sample size for Egyptian Arabic separately from every other Arabic variety in the project — Egyptian is the easiest Arabic to collect and will dominate by default, and the buyer needs the real distribution to know what the model is learning.

  • Arabic (Gulf) Multilingual Speech

    Lock the label set before collection and state where expatriate-language speech goes — its own language label, or excluded from delivery — since Urdu, Malayalam, and Tagalog arrive in the same recordings as Gulf Arabic.

  • Kannada Multilingual Speech

    State whether Tulu, Konkani, and Urdu material inside Kannada recordings gets its own language label or is treated as borrowing — coastal and northern Karnataka speakers mix them in as a matter of course, and an unlabelled mix leaves the batch's composition unexplainable.

  • Punjabi Multilingual Speech

    Split the Punjabi label into a Gurmukhi-side and a Shahmukhi-side sub-label, with explicit decision rules against Urdu for the Shahmukhi side — a merged Punjabi label is not recoverable after annotation.

  • Burmese Multilingual Speech

    State whether Rakhine is carried as its own label or folded into Burmese, and define how the Burmese-derived scripts of Mon and Shan are told apart before automatic language identification is applied.

  • Amharic Multilingual Speech

    The delivery must state character encoding and normalization status per language, and supply the decision rules used to separate Amharic from Tigrinya recordings inside the same project.

  • Ukrainian Multilingual Speech

    Lock the label scheme up front — split labels with a dedicated mixed category, or a merged label with sub-tags — and supply decision rules with borderline examples for the Ukrainian-Russian boundary.

  • Swahili Multilingual Speech

    Language label granularity has to be locked at project start — utterance-level, not speaker-level — and the Arabic loanword layer has to be defined as Swahili in writing, with a decision list, or teams will label the same word differently.

  • Kurdish Multilingual Speech

    Hours, speaker counts, and recording conditions are listed per branch, and the two branches are annotated and delivered by separate teams with separate quality reports, never merged into one figure.

  • Hausa Multilingual Speech

    Provide decision rules for embedded Arabic formulaic phrases and a sub-label separating first-language from second-language Hausa; without the sub-label the batch's phonology is bimodal for reasons the data cannot explain.

  • Uzbek Multilingual Speech

    Define the label granularity before collection — Uzbekistan, Afghanistan, and the diaspora communities kept separate — and record each speaker's country and script of schooling item by item.

About Multilingual Speech →

Noisy Speech

26 languages

  • Hindi Noisy Speech

    Noise type must be annotated item by item, and it has to include Indian local scenes — street, market, auto-rickshaw. A generic noise library is not a substitute.

  • Arabic (MSA) Noisy Speech

    State whether speakers were asked to hold a natural speaking rate during recording; if they were told to raise their voice, record that in the metadata.

  • Indonesian Noisy Speech

    Annotate the noise scene item by item, and state whether language-component annotation is also done on the noisy recordings.

  • Thai Noisy Speech

    Set SNR tiers to a tonal-language standard — one tier cleaner than the general benchmark is a reasonable default — and annotate the measured value for every clip.

  • Turkish Noisy Speech

    Annotate word-final clarity or provide a per-word intelligibility score. A sentence-level SNR figure on its own is not enough.

  • Vietnamese Noisy Speech

    Annotate the spectral characteristics of each noise type, and prioritize local scenes such as motorbike traffic.

  • Filipino Noisy Speech

    The annotation guideline has to specify noise annotation and code-switching annotation together. Neither one substitutes for the other.

  • Persian Noisy Speech

    Persian script omits short vowels, so a word heard through noise cannot be recovered from the writing system — the annotation needs an explicit uncertainty convention rather than a spelling guess.

  • Tamil Noisy Speech

    Whether the annotator may collapse Tamil's nasal series to a single 'n' — the language has six contrastive nasals, and a degraded recording often does not carry the place cue.

  • Telugu Noisy Speech

    Annotators have to pass a dialect identification test, and the test records should be included in the delivery.

  • Malayalam Noisy Speech

    Assess the timeline separately rather than at the rate of other languages, and expect lower output per hour of audio.

  • Nepali Noisy Speech

    The noise scene has to be reproducible — location, time of day, and equipment recorded item by item — or the data cannot be replicated.

  • Sinhala Noisy Speech

    Measure the loanword share in noisy conditions and compare it against a quiet-environment baseline, with the difference explained.

  • Bangla Noisy Speech

    Record the recording distance and whether a windscreen was used, item by item. Without them the proximity effect cannot be accounted for.

  • Khmer Noisy Speech

    Khmer carries its lexical load in consonants rather than in tone, so noise attacks exactly the distinction the transcript depends on — fix an uncertainty convention instead of letting annotators guess.

  • Arabic (Egyptian) Noisy Speech

    Annotate the noise scene item by item (Cairo traffic, market, horn density) and record the measured signal-to-noise ratio per file, with tiers set one step cleaner than a non-reducing variety would need — the glottal-stop reduction leaves no consonant to recover.

  • Arabic (Gulf) Noisy Speech

    Annotate noise type and signal-to-noise ratio per file, and state whether the overlapping speech of multi-speaker settings counts as noise or is preserved as legitimate content.

  • Kannada Noisy Speech

    Classify each file's noise by type and state whether it contains intelligible speech; where it does, annotate the competing speech or exclude the file — one SNR figure does not distinguish traffic from a crowd talking.

  • Punjabi Noisy Speech

    Fix a flag for tone-uncertain words — mark the uncertainty rather than letting annotators pick a spelling — and annotate the noise scene and side of collection on every clip.

  • Burmese Noisy Speech

    Set the SNR tiers so that creaky-tone and breathy-tone items get a cleaner tier than the general benchmark, and annotate the measured SNR for every clip rather than a batch-level figure.

  • Amharic Noisy Speech

    Log the noise scene and the measured signal-to-noise ratio per file, and fix an uncertainty convention so annotators mark unintelligible stretches rather than reconstructing them from context.

  • Ukrainian Noisy Speech

    Annotate a measured signal-to-noise ratio per file plus a flag for whether the stressed syllable remains audible; a sentence-level SNR figure alone does not capture the risk.

  • Swahili Noisy Speech

    The noise scene has to be annotated item by item and be reproducible (location, time of day, device), with the measured SNR value attached per clip — a generic "noisy" label can be neither checked nor replicated.

  • Kurdish Noisy Speech

    Every file carries an SNR band, a noise type, and a location category; because accessible environments cluster in homes and community spaces, stating which categories the batch actually covers is part of the delivery.

  • Hausa Noisy Speech

    Set SNR tiers tighter than the general benchmark, annotate the measured value for every clip, and note whether tone and length remain audible at that level.

  • Uzbek Noisy Speech

    Annotate the noise scene item by item, and state per file whether numerals and measurements were spoken in Uzbek or Russian — after collection that is not recoverable from the audio.

About Noisy Speech →

Code-Switching Speech

26 languages

  • Hindi Code-Switching Speech

    How the Hindi portion is written inside a Hinglish utterance — Devanagari, Roman, or mixed — because that choice decides whether the transcript is reusable in an English-trained pipeline.

  • Arabic (MSA) Code-Switching Speech

    Distinguish dialect-to-MSA switching from Arabic-to-foreign-language switching as two categories, annotated separately and never merged.

  • Indonesian Code-Switching Speech

    The annotation scheme must support switching among three or more languages, with a language label on every word position. Binary annotation will not do.

  • Thai Code-Switching Speech

    The nativized-loanword list has to be settled before annotation starts, not during — Thai nativization is gradual enough that annotators will disagree on the borderline cases.

  • Turkish Code-Switching Speech

    Annotate each speaker's age and language background — they are what explains the differences in word choice.

  • Vietnamese Code-Switching Speech

    Make the tone annotation for nativized English words explicit — annotate what was actually said — and state whether these words are counted separately.

  • Filipino Code-Switching Speech

    Annotate the level of each switch — word, phrase, or sentence — separately. One blanket switch marker is not enough.

  • Persian Code-Switching Speech

    How far the Arabic loanword layer counts as Persian — it is fully nativized in speech, and treating it as switching would label ordinary Persian conversation as code-switching.

  • Tamil Code-Switching Speech

    The annotation scheme has to express both language switching and register difference. Labeling language alone is not enough.

  • Telugu Code-Switching Speech

    Separate absorbed Urdu loanwords from live switching, and provide a decision list.

  • Malayalam Code-Switching Speech

    Annotators have to pass an agreement test — several people annotating the same batch, with the agreement rate calculated. Batches that fall short get reworked.

  • Nepali Code-Switching Speech

    Annotate the language of education for each speaker, and deliver stratified by switching density.

  • Sinhala Code-Switching Speech

    English components stay in and get annotated, with a share statistic provided. Cleaning to pure Sinhala is not acceptable.

  • Bangla Code-Switching Speech

    Whether the English words are written in Bangla script or in English spelling — both conventions are in use in Bangla transcription, and the choice decides whether the corpus can be compared against English data.

  • Khmer Code-Switching Speech

    Settle one convention for writing English loanwords — transliteration or Latin letters kept as is — and put examples in the guideline.

  • Arabic (Egyptian) Code-Switching Speech

    The guideline has to define the three layers separately and supply a decision list for the nativized French and Italian loanwords that are never switch points — annotators will disagree on the borderline cases without one.

  • Arabic (Gulf) Code-Switching Speech

    Fix the script convention for English insertions (Latin inside Arabic text, or Arabic transliteration), the switch-point granularity, and a separate tag for Arabic-to-Arabic variety switching.

  • Kannada Code-Switching Speech

    State whether annotation marks switch spans or tags language per word — with English material carrying Kannada inflections, per-word tagging scatters single-word tags instead of spans, and the two outputs train different models.

  • Punjabi Code-Switching Speech

    The annotation scheme carries the speaker's side as a required dimension, with a separate switch-type inventory and decision list for each side — one merged guideline produces internally contradictory labels.

  • Burmese Code-Switching Speech

    Define three annotation classes — historical loanword, live single-word switch, sentence-level insertion — with a decision list and worked examples, and never merge them into one switch label.

  • Amharic Code-Switching Speech

    One writing convention for English insertions has to be locked with worked examples, either transliterated into Ethiopic characters or Latin letters kept as they are, and applied by every annotator on the project.

  • Ukrainian Code-Switching Speech

    Annotate surzhyk on two axes — dominant language and mixing-degree band — instead of forcing a binary switch marker, and state how English content on top of that is tagged.

  • Swahili Code-Switching Speech

    State which mixing pattern the batch targets (Swahili-English / Swahili with a local language / Sheng) and annotate switch points at word level with a language label per token — the three patterns cannot share one annotation scheme.

  • Kurdish Code-Switching Speech

    Whether switch points are marked at word level and whether the Kurdish portion is tagged by branch are both fixed in the guideline, because those two decisions set the annotation workload and the tooling.

  • Hausa Code-Switching Speech

    Fix the nativized-loanword list before annotation begins (Arabic religious phrases in, not counted as switches), and annotate the tone and length actually assigned to English words rather than a dictionary form.

  • Uzbek Code-Switching Speech

    Fix the nativized-loanword list before annotation begins and record the speaker's age band — the same recording yields different switch counts depending on who annotates it and how the list was drawn.

About Code-Switching Speech →

Children Speech

26 languages

  • Hindi Children Speech

    Age-band grouping (3–5, 6–8, 9–12 years, for example) and guardian authorization documents are both mandatory attachments. Without either one, nothing ships.

  • Arabic (MSA) Children Speech

    State the family dialect background of each child speaker, and record an assessment of their proficiency in MSA.

  • Indonesian Children Speech

    The child's age at first exposure to Indonesian has to be recorded — children who learned it at school carry a different sound system from those who heard it at home.

  • Thai Children Speech

    Keep age bands to a span of 3 years or less, and attach a note on tone production for each band.

  • Turkish Children Speech

    Record age in months, not just an age band, and log the sampling date for every child.

  • Vietnamese Children Speech

    Annotate the child's dialect region — north, central, or south — and an exact age. Both are required.

  • Filipino Children Speech

    Annotate the child's language environment — home language and school language — and the share of mixed speech.

  • Persian Children Speech

    Collection must use spoken-form elicitation rather than reading aloud — Persian's written and colloquial forms diverge enough that a child reading written text produces a register no child actually speaks.

  • Tamil Children Speech

    The annotation guideline has to separate valid colloquial forms from pronunciation deviations as two categories, with examples for each.

  • Telugu Children Speech

    Annotate the child's home dialect region, and allow that dialect's valid forms into the transcription.

  • Malayalam Children Speech

    Assess the schedule separately and allow a full recruitment period. Estimating at adult-collection speed is not acceptable.

  • Nepali Children Speech

    State the collection venue and how it was organized, and provide records of the guardian informed consent process.

  • Sinhala Children Speech

    Annotate where English loanwords appear, and describe each child's level of exposure to English.

  • Bangla Children Speech

    The age band has to be fixed and reported per speaker — Bangla children's vocabulary and pronunciation shift sharply between the 3–5 and 9–12 bands, and a single 'child' label hides the split.

  • Khmer Children Speech

    No task may depend on reading or writing — Khmer script takes years to master, and children's command of it varies far too widely to use as a screening criterion.

  • Arabic (Egyptian) Children Speech

    Annotate each child's region (Cairo / Alexandria / Upper Egypt) and state the family dialect — the qaf is realized differently in the two regions, and the annotation standard has to accept both rather than treating one as a defect.

  • Arabic (Gulf) Children Speech

    Record each child's family dialect, home language mix, and school language, and confirm that the guardian consent covers the child's voice being used for model training.

  • Kannada Children Speech

    Record each child's medium of instruction and home language alongside the age band — urban Karnataka's English-medium shift makes the school language a stronger predictor of a child's Kannada than age, and the two groups need separate rules for cluster simplification.

  • Punjabi Children Speech

    Keep tasks free of reading and writing, or match the task script to the child's schooling; annotate each child's home language, schooling language, and side — a single "Punjabi child" label hides all three profiles.

  • Burmese Children Speech

    State that children's colloquial forms are transcribed as spoken and never normalized toward literary Burmese, and attach a list of the common forms annotators most often try to correct.

  • Amharic Children Speech

    No task may depend on reading beyond the child's grade level, and the child's language of instruction at school has to be recorded next to the home language.

  • Ukrainian Children Speech

    Record the family's home language alongside the child's age in months, and attach the guardian consent documentation to every delivery. Age bands alone are not enough.

  • Swahili Children Speech

    Age band (narrow, in years), the child's home language, and guardian consent documents are all mandatory attachments per record — without the home language there is no way to separate first-language transfer from a developmental pattern.

  • Kurdish Children Speech

    Age bands and signed guardian consent are delivered with every batch, and each recording notes the child's school language, since it predicts where the substitutions fall.

  • Hausa Children Speech

    Record the child's age in months and the home language (whether Hausa is a first language), and keep age bands narrow; a wide band averages together children at visibly different stages of tone acquisition.

  • Uzbek Children Speech

    Record each child's age in months, the script used at home, and the language of schooling — the two-script environment is the main source of variation in this data.

About Children Speech →

Audiobook

26 languages

  • Hindi Audiobook

    Whether timbre has to stay consistent across chapters — the same speaker recording over multiple days — is what sets the schedule and the cost.

  • Arabic (MSA) Audiobook

    Provide a fully vocalized version of the reading text, and state the rule for choosing among words that have more than one accepted reading.

  • Indonesian Audiobook

    Build a loanword pronunciation table and include it with the delivery, so reading consistency can be explained and checked.

  • Thai Audiobook

    Provide the word-segmented, annotated text as a delivery attachment, and state the phrase-break convention.

  • Turkish Audiobook

    How dialogue is attributed — Turkish drops pronouns freely, so a transcript of a dialogue-heavy book does not say who is speaking and the annotation has to take it from the narration.

  • Vietnamese Audiobook

    Annotate the reader's dialect region, and decide in advance whether readers from different dialect regions can be mixed. Our recommendation is not to mix them.

  • Filipino Audiobook

    Annotate language by paragraph, and state the reading convention for English passages — which pronunciation standard, and at what speaking rate.

  • Persian Audiobook

    Whether narration follows the written form or is allowed to colloquialize — literary Persian narration produces a register that almost never occurs in spontaneous speech, which limits what the corpus can train.

  • Tamil Audiobook

    Whether the narrator reads literary Tamil or a spoken variant — Tamil diglossia is extreme enough that a listener trained on one cannot follow the other, and mixing them within one book makes the corpus unusable.

  • Telugu Audiobook

    Provide a pronunciation reference for the Sanskritized vocabulary, and run a test read to check it before the real recording starts.

  • Malayalam Audiobook

    Lock the number of readers and their available hours in advance, and run a test recording to assess long-form consistency.

  • Nepali Audiobook

    Confirm and document the rights status of every text item by item, and write the reading convention for honorific forms into the guide.

  • Sinhala Audiobook

    Whether narration uses literary or colloquial Sinhala — the two differ in verb forms and pronouns throughout, and a book that mixes them is usable for neither.

  • Bangla Audiobook

    Text and reader must be matched by region. Annotate register and run QC separately for each register.

  • Khmer Audiobook

    State whether the transcription uses orthography or phonemic notation, and provide a correspondence table as a delivery attachment.

  • Arabic (Egyptian) Audiobook

    Whether the reading text is written in a fixed Egyptian orthography and which convention governs it — with no standard spelling for the dialect, the text itself is the specification, and it has to ship with the delivery so the audio can be checked against it.

  • Arabic (Gulf) Audiobook

    State whether the narration text is standard or Gulf dialect and whether Gulf coloring in a standard reading is accepted; if dialect, attach the spelling convention used in the text.

  • Kannada Audiobook

    Run cluster-level audio-versus-text checks on a sample of every chapter, with a stated re-record trigger — word-level comparison will not catch a misread consonant inside a fused cluster, which is this script's characteristic failure mode.

  • Punjabi Audiobook

    Match the manuscript script and the narrator's side as a single locked pair, and run tone-consistency QC at chapter level — the text offers no reference for it, so the check is audio-only.

  • Burmese Audiobook

    Deliver the narration text with a pronunciation reference for the Pali-derived vocabulary, and run a test read of a full chapter before the recording schedule is committed.

  • Amharic Audiobook

    Supply one pronunciation policy for the whole book covering recurring names and terms, and state whether the same narrator records every chapter, since changing narrators mid-book breaks timbre continuity.

  • Ukrainian Audiobook

    Deliver the stress-marked narration text with the audio, and state the re-record policy for chapters where the reader's stress choices diverge from the marked text.

  • Swahili Audiobook

    Whether narration stays in standard Swahili throughout or is allowed to colloquialize, and whether timbre must hold across chapters recorded on different days — together those two set the schedule and the cost.

  • Kurdish Audiobook

    The narration text's orthographic standard and script are fixed before recording, and the delivery states whether the narrator's own dialect forms were corrected in the session or left in place.

  • Hausa Audiobook

    Provide a tone and length marked reference text, run a test read to check the reader against it, and lock the reader's variety to the text's, which is Kano-based; a reader from another variety will not realize it the same way.

  • Uzbek Audiobook

    State which script the narrator reads from and how a converted text was verified — a machine-converted text handed straight to a reader produces hesitation and mispronunciation that no audio QC can repair.

About Audiobook →

Short Video Speech

26 languages

  • Hindi Short Video Speech

    Whether background music stays or is removed directly determines which use cases the data can serve. Settle it first.

  • Arabic (MSA) Short Video Speech

    Write out the register screening criteria for content sources, and annotate the register distribution of what was actually collected in the delivery.

  • Indonesian Short Video Speech

    State the background music handling — kept, separated, or removed — and whether the isolated dry vocal is delivered. Both need to be explicit.

  • Thai Short Video Speech

    Decide in advance whether word segmentation is annotated and whether particles are annotated separately.

  • Turkish Short Video Speech

    Phoneme annotation has to allow emotion-driven shifts in phonetic value, with a note on the tolerated range.

  • Vietnamese Short Video Speech

    Annotate tones as actually produced, without normalizing to standard pitch values, and note how speaking rate affects pitch values.

  • Filipino Short Video Speech

    Annotate language by sentence and report the share of English content. Do not process it as pure Filipino.

  • Persian Short Video Speech

    Whether the captions the creator burned into the video are used as a transcription reference — informal Persian captions are written loosely and often disagree with what is actually said.

  • Tamil Short Video Speech

    Transcribe in the colloquial form. Annotate English phrases separately rather than forcing them into Tamil.

  • Telugu Short Video Speech

    How to treat clips that begin or end mid-word — Telugu's agglutinative suffixes mean a cut fragment carries a partial word, and transcribing it as a whole word teaches the model wrong morphology.

  • Malayalam Short Video Speech

    Lock the sampling standard in advance — topic, duration, creator type — and estimate the workload from it.

  • Nepali Short Video Speech

    State the content source range and topic coverage, and report the topic distribution actually collected.

  • Sinhala Short Video Speech

    Transcribe in standard Sinhala script, not Latin transliteration. If the content sources contain Latin-script text, state how it is handled.

  • Bangla Short Video Speech

    How the burned-in captions are treated — Bangla short-video creators often caption in Romanized Bangla rather than Bangla script, and the transcript convention has to match whichever the buyer's pipeline expects.

  • Khmer Short Video Speech

    How sped-up and music-bedded segments are treated — whether to transcribe what the creator spoke at natural speed or what is audible in the published cut, and how to tag segments that are unrecoverable.

  • Arabic (Egyptian) Short Video Speech

    How burned-in Franco-Arabic captions are handled — Egyptian creators caption in Latin letters with digits (3, 7, 2) standing in for Arabic sounds, so the delivery has to state whether the transcript ships in Arabic script or in the creator's own spelling.

  • Arabic (Gulf) Short Video Speech

    Fix the transcription convention for Quranic quotations and religious formulas — fixed standard orthography, never dialect spelling — as a rule separate from the surrounding colloquial text.

  • Kannada Short Video Speech

    Stratify sampling by creator niche and report language composition per stratum — niche, not topic, is what predicts English density in Kannada short video, and an unstratified sample produces an unexplainable two-peak distribution.

  • Punjabi Short Video Speech

    State the caption script inventory of the source pool and lock the delivery transcript to one script; Roman captions are used as a cross-check only and never as the transcription source.

  • Burmese Short Video Speech

    Annotate the tones as produced, without restoring the citation form, and state explicitly that burned-in captions are not used as a transcription reference.

  • Amharic Short Video Speech

    Transcribe in Ethiopic script regardless of how the creator captioned the clip, and state how Latin-script captions and on-screen text were handled when they diverged from the audio.

  • Ukrainian Short Video Speech

    Annotate language per sentence, and add a flag for clips where the burned-in caption language differs from the spoken language, so the two are never conflated.

  • Swahili Short Video Speech

    Whether the background music is kept or removed, and whether Sheng and English insertions get their own language tags — the first decides what the data can train, the second decides whether the language makeup can be measured at all.

  • Kurdish Short Video Speech

    Each clip carries a branch tag and the creator's country of residence, and the delivery states whether background music was removed or retained, since that changes what the data can train.

  • Hausa Short Video Speech

    Transcribe from the audio rather than from the burned-in captions, and annotate the Arabic and English segments separately instead of folding them into the Hausa text.

  • Uzbek Short Video Speech

    Transcribe in standard Latin orthography rather than copying caption spellings, and state in the delivery notes how Cyrillic-script captions and comments are handled.

About Short Video Speech →

Live Stream Speech

26 languages

  • Hindi Live Stream Speech

    How long a single session runs, and whether delivery is cut into fixed segments (10-second clips, for example) — both have to be settled first.

  • Arabic (MSA) Live Stream Speech

    Annotate the streamer's dialect and the audience's origin per session, and deliver grouped by dialect.

  • Indonesian Live Stream Speech

    Sampling must exclude product-selling sessions with highly repetitive patter, or deduplicate the repeated content and state what proportion was removed.

  • Thai Live Stream Speech

    Whether overlap annotation is done at all, and whether overlapping segments are cut out separately — both settled in advance.

  • Turkish Live Stream Speech

    The sensitive-content filter rules have to be set in advance, and the delivery notes must state what proportion was filtered and on what criteria.

  • Vietnamese Live Stream Speech

    The streamer's dialect and the dialects of the interacting viewers are annotated separately — one overall label is not enough.

  • Filipino Live Stream Speech

    Background-music handling and language annotation are defined as two separate items — neither stands in for the other.

  • Persian Live Stream Speech

    Compliance clearance for source material has to be documented; the transcription standard is specified as spoken form.

  • Tamil Live Stream Speech

    Annotators must complete spoken-transcription training and pass a test, with the test records delivered alongside the data.

  • Telugu Live Stream Speech

    Whether audience chat and donation alerts read aloud by the host are transcribed — they are a large share of live-stream audio and are not the host's own speech.

  • Malayalam Live Stream Speech

    Quote and schedule this language on its own actual difficulty — do not carry over the per-hour rate used for other languages.

  • Nepali Live Stream Speech

    State the range of source material and the time span it covers, and report topic coverage statistics.

  • Sinhala Live Stream Speech

    Transcription uses standard Sinhala script throughout; state how Latin-script content is handled.

  • Bangla Live Stream Speech

    Whether long sessions are delivered whole or cut into fixed-length segments, and how the host's read-aloud of viewer comments is marked in the transcript.

  • Khmer Live Stream Speech

    Measure and deliver the register distribution, stratified by register; state the range of source material.

  • Arabic (Egyptian) Live Stream Speech

    The sensitive-content rules have to name the pervasive religious and blessing expressions as ordinary speech rather than filter targets, and the delivery must state what share was filtered and on which criteria — otherwise the cleanup silently changes the register of the corpus.

  • Arabic (Gulf) Live Stream Speech

    Label every session by type (commerce, religious, general) before scoping, and fix the number-normalization rule and the quotation-transcription rule as two separate conventions.

  • Kannada Live Stream Speech

    Mark host-read viewer comments as a separate segment type and transcribe the host's spoken rendering as speech — do not align the transcript against the typed comment feed, since romanized Kannada comment text and the host's rendering are never the same string.

  • Punjabi Live Stream Speech

    Annotate per session the audience origin and the language composition, and fix how Roman-script viewer names and messages are represented in the transcript — the reading-aloud convention cannot be recovered afterwards.

  • Burmese Live Stream Speech

    Set the sensitive-content filter rules in advance and report the filtered proportion and the criteria in the delivery notes; confirm the team and data handling arrangement before the schedule is fixed.

  • Amharic Live Stream Speech

    Fix the session length and the segmentation rule first, then state how proper names were transcribed: whether a name list was agreed with the streamer beforehand or written as heard.

  • Ukrainian Live Stream Speech

    Record each session's date and its observed language composition, and state the collection's date range, because the archive spans the point where the streamer changed language.

  • Swahili Live Stream Speech

    Session length and whether the recording is delivered whole or cut into fixed segments have to be settled first, along with the rule for marking read-aloud viewer comments so they are not attributed to the host.

  • Kurdish Live Stream Speech

    Session length and the fixed segment length for delivery are set in advance, and the content-filtering rule is written per source country rather than as one blanket rule across the batch.

  • Hausa Live Stream Speech

    Annotate the host's Hausa variety and the first language of each interacting speaker separately, and state whether read-aloud viewer comments are transcribed at all; they are a large share of the audio and are not the host's own speech.

  • Uzbek Live Stream Speech

    Annotate read-aloud viewer comments separately from the host's own speech, and state the script convention for the transcript — the two speech types cannot share one label.

About Live Stream Speech →

Voice Assistant

26 languages

  • Hindi Voice Assistant

    Whether the intent schema comes from you or is designed by us is the line the quote falls on, and it has to be settled first.

  • Arabic (MSA) Voice Assistant

    The intent schema must state which dialects it covers and include a mapping table from dialect expression to intent.

  • Indonesian Voice Assistant

    The intent schema must normalize across languages and provide the mapping between the expressions used in each.

  • Thai Voice Assistant

    State whether polite variants are merged into one intent, and report how the variants are distributed.

  • Turkish Voice Assistant

    The slot-extraction rules must state how suffix variants are handled, with an example set of variants provided.

  • Vietnamese Voice Assistant

    State whether pronoun variants are normalized, and report how users are distributed across age groups.

  • Filipino Voice Assistant

    The intent schema must cover Taglish expression, with an example set of mixed-language utterances.

  • Persian Voice Assistant

    Intent definitions are written in spoken form, with a spoken-to-written correspondence table provided.

  • Tamil Voice Assistant

    Intent definitions are written in spoken form; annotators must pass a spoken-transcription test.

  • Telugu Voice Assistant

    The intent schema must normalize across dialects and state the sample size for each.

  • Malayalam Voice Assistant

    Whether slot annotation is done by the same annotators as the transcription — Malayalam intent-plus-slot work needs stronger language ability than transcription, and the two pools are not the same size.

  • Nepali Voice Assistant

    Which honorific register the assistant's responses target — Nepali has three levels of honorific verb forms, and an intent schema written without one produces responses that are wrong for the user.

  • Sinhala Voice Assistant

    English command words are brought into the intent schema; the urban-rural sample distribution is measured and stated.

  • Bangla Voice Assistant

    The intent schema has to be validated against each market's phrasing separately, stating sample size and coverage for both.

  • Khmer Voice Assistant

    State whether register variants are normalized, and provide a register correspondence word list.

  • Arabic (Egyptian) Voice Assistant

    The slot-extraction rules have to state how entity values containing a qaf are matched — the spoken glottal stop and the written standard letter are the same slot value, and a matcher built on the written form will miss every one of them.

  • Arabic (Gulf) Voice Assistant

    Attach a Gulf pronunciation lexicon to the command list, marking every written form pronounced with /ɡ/, /tʃ/, or a regional ج, and state a separate exact-match rule for chapter and proper names.

  • Kannada Voice Assistant

    Define slot-value handling for English fillers inside Kannada utterances, with examples per slot type, and fix whether slot values are written in Latin or in Kannada script — the choice decides whether the corpus merges with an English pipeline.

  • Punjabi Voice Assistant

    Deliver the intent schema in both scripts with the same intent set, and attach a per-side example set for every intent — annotators cannot work from a schema written in the other side's script.

  • Burmese Voice Assistant

    Build the intent schema in spoken form with a pronunciation key for slot values, and state that personal names are matched by pronunciation because their spelling is not stable.

  • Amharic Voice Assistant

    The intent schema must state the clock and calendar convention it normalizes to, and whether honorific and plain address forms are merged into one intent or kept apart with their own examples.

  • Ukrainian Voice Assistant

    The intent schema states whether Russian-derived habitual command forms are covered, and how mixed-language utterances are normalized onto one intent, with example sets attached.

  • Swahili Voice Assistant

    Whether the intent schema is supplied by the buyer or designed from scratch, and whether slots are annotated with the language they surface in — Swahili sentences routinely carry English slot values, and a schema that ignores this loses the slot at extraction time.

  • Kurdish Voice Assistant

    The intent schema is defined separately for each branch, and the delivery states whether it came from the buyer or was designed for the project, with slot-value conventions written out.

  • Hausa Voice Assistant

    Decide whether the Hausa and English phrasings of one intent form a single class or separate classes, and whether intent definitions carry tone and length marking so annotators can match what was actually said.

  • Uzbek Voice Assistant

    State whether Russian and mixed-language commands are in scope for the intent schema, and which lexical variant the example set treats as canonical for each intent.

About Voice Assistant →

Elderly Voice

26 languages

  • Hindi Elderly Voice

    The age stratification and whether speakers with speech disorders are included must be stated up front, not added afterward.

  • Arabic (MSA) Elderly Voice

    Annotate dialect background record by record; state up front whether speakers with speech disorders are included.

  • Indonesian Elderly Voice

    Whether the batch includes speakers whose Indonesian is a second language at all — older rural speakers may be more fluent in Javanese, and a 'clear Indonesian' requirement has to be stated explicitly.

  • Thai Elderly Voice

    Age stratification has to be fine-grained (60–69 / 70–79 / 80+, for example), with a note on tone production for each band.

  • Turkish Elderly Voice

    Record each speaker's era of schooling and vocabulary preference, to explain the differences in vocabulary distribution.

  • Vietnamese Elderly Voice

    Annotate dialect region record by record; north and south do not share a batch, and state the age distribution within each dialect region.

  • Filipino Elderly Voice

    Measure the share of English content and state it against comparable data from younger groups.

  • Persian Elderly Voice

    Record each speaker's register tendency, and describe the register distribution in the delivery notes.

  • Tamil Elderly Voice

    Record each speaker's colloquial-versus-written tendency, and state it against a comparison with younger groups.

  • Telugu Elderly Voice

    Assess the recruitment timeline separately; if the sample falls short, say so in the delivery notes rather than padding it with another age group.

  • Malayalam Elderly Voice

    Deliver in batches by age band, with speaking-rate metrics measured separately for each.

  • Nepali Elderly Voice

    Record the recording device and environment for every item; state the urban-rural sample distribution.

  • Sinhala Elderly Voice

    Measure the share of English content separately for each age band; the age stratification has to be fine-grained.

  • Bangla Elderly Voice

    Deliver grouped by region, and describe the use of old-fashioned vocabulary on each side.

  • Khmer Elderly Voice

    Record register usage; deliver in batches by age band and describe the register distribution for each.

  • Arabic (Egyptian) Elderly Voice

    State the age stratification precisely and record whether each speaker retains final vowels and full clusters — the elision that defines Cairo colloquial is age-graded, and a batch of younger speakers will not contain the unreduced forms at all.

  • Arabic (Gulf) Elderly Voice

    Record whether each speaker can read a prepared script at all and choose the elicitation mode accordingly, and supply a glossary of pre-oil-era vocabulary to the annotation team.

  • Kannada Elderly Voice

    Annotate each speaker's schooling medium and era along with the age band — the elderly cohort's lower English mixing is a cohort effect rather than an age effect, and without the schooling annotation the two cannot be separated.

  • Punjabi Elderly Voice

    Record each speaker's birth region and, where relevant, migration history, alongside a fine-grained age band — a modern regional label alone does not describe what a partition-generation speaker actually carries.

  • Burmese Elderly Voice

    Use age bands no wider than a decade, record the recording environment and device for every item, and state whether the batch represents the older sound system or current usage.

  • Amharic Elderly Voice

    State the age stratification and record for every speaker whether Amharic is a first language or a second, along with the schooling language, since that split explains most of the vocabulary variation.

  • Ukrainian Elderly Voice

    Annotate each speaker's language of schooling and everyday dominant language before recording, and set the acceptance rule for Russian-dominant speakers in writing.

  • Swahili Elderly Voice

    Age stratification has to be fine-grained (60-69 / 70-79 / 80+, for example), and each speaker's first language and years of formal schooling recorded — state up front whether speakers with speech or hearing conditions are included.

  • Kurdish Elderly Voice

    Age bands are set in advance, and each speaker is screened for literacy in the target script; if reading tasks are impossible, the plan switches to prompted spontaneous speech rather than dropping the speaker.

  • Hausa Elderly Voice

    Record each speaker's literacy script (Boko, Ajami, or both) alongside the age band, and design elicitation around it; a Boko reading task quietly excludes the speakers whose Ajami literacy is the point.

  • Uzbek Elderly Voice

    Record each speaker's script of schooling and age band, and set elicitation materials in the script the speaker actually reads — Cyrillic for the oldest band, Latin for the youngest.

About Elderly Voice →

Accented English

26 languages

  • Hindi Accented English

    Accent grouping has to break down by native-language background (Hindi / Tamil / Bangla / Gujarati and so on), not just by country.

  • Arabic (MSA) Accented English

    Group by country of origin (Egypt / Saudi Arabia / the UAE / Lebanon and so on), and record each speaker's English proficiency.

  • Indonesian Accented English

    Annotate both accent source and English proficiency level, each with stated criteria for how it was assessed.

  • Thai Accented English

    Record each speaker's English proficiency; the accent annotation must state whether syllable-structure-related systematic deviations are distinguished.

  • Turkish Accented English

    Which first-language features the batch must actually contain — Turkish lacks /θ/, /ð/ and /w/ and devoices final obstruents, and a batch of highly proficient speakers will not carry any of them.

  • Vietnamese Accented English

    Whether the batch separates Northern, Central and Southern speakers — the three transfer different features into English, and Southern speakers drop final consonants far more heavily.

  • Filipino Accented English

    Annotate English proficiency level, and state where each speaker acquired English (home / school / work).

  • Persian Accented English

    Annotate the speaker's place of residence and English acquisition environment; those two account for most of the accent variation.

  • Tamil Accented English

    Group by the speaker's country or region — Tamil background alone is not a usable label.

  • Telugu Accented English

    Accent labels have to break down to the native-language level, with stated criteria for distinguishing them from neighboring native-language accents.

  • Malayalam Accented English

    Set the sample-size target from actual recruitment capacity, and state it honestly in the delivery notes.

  • Nepali Accented English

    Record each speaker's English learning path (whether Hindi mediated it), which explains where the accent variation comes from.

  • Sinhala Accented English

    Annotate accent source and proficiency level together, and state the urban-rural sample distribution.

  • Bangla Accented English

    Group by the speaker's region, and record the environment where English was acquired.

  • Khmer Accented English

    Annotate English proficiency level with the assessment method stated; the sample has to span different proficiency levels.

  • Arabic (Egyptian) Accented English

    Group by the speaker's school language track (English-medium / French-medium / public) and record English proficiency against stated criteria — the tracks transfer different vowel systems, and a single Egyptian label flattens exactly the variation a robustness set needs.

  • Arabic (Gulf) Accented English

    Record each speaker's school type, the first language of their English teachers, and their Arabic variety, instead of one country-level accent label.

  • Kannada Accented English

    Record each speaker's English acquisition setting and current usage, and state the recruitment channel — a batch drawn from Bangalore's professional population may carry almost none of the transfer features the label implies.

  • Punjabi Accented English

    Group by acquisition environment (Indian Punjab, Pakistani Punjab, UK diaspora, Canada diaspora) and record whether the speaker was raised in the home market or abroad — the two are not interchangeable for accent robustness work.

  • Burmese Accented English

    Group speakers by where and how English was acquired (state school / English-medium / abroad) and record a proficiency assessment for each, since that split accounts for more variation than the accent label alone.

  • Amharic Accented English

    Record proficiency level with the assessment method stated, and record where English was acquired: home, private school, university, or work. Within Ethiopia the acquisition path explains more accent variation than geography does.

  • Ukrainian Accented English

    Group by first language and record whether Russian was also an active language for the speaker; separate domestic speakers from diaspora speakers in the delivery.

  • Swahili Accented English

    Group by first-language background (Swahili / Kikuyu / Luo / Luhya and so on) and record country of origin separately — the two schemes answer different questions and neither substitutes for the other.

  • Kurdish Accented English

    Record each speaker's country of origin and the language that mediated their English learning; accent labels group by that route, and the delivery states how many speakers each route contributed.

  • Hausa Accented English

    Accent labels break down to the speaker's first language (Hausa separately from Yoruba, Igbo, and others), and record English proficiency and where it was acquired, which together explain most of the spread within the label.

  • Uzbek Accented English

    Record each speaker's language of instruction and whether English was learned through Russian; without that, the two transfer layers cannot be told apart in the batch.

About Accented English →

Speech Translation

26 languages

  • Hindi Speech Translation

    Whether the translation is human or machine translation with human post-editing is an order-of-magnitude difference in cost and quality, and it has to be settled first.

  • Arabic (MSA) Speech Translation

    State the source form (standard or dialect) and the target register separately; a general Arabic is not accepted.

  • Indonesian Speech Translation

    Which of the three layers in the audio gets translated — Indonesian, the regional language, and English each need their own rule, and one blanket instruction produces inconsistent output.

  • Thai Speech Translation

    The source transcription is delivered with a word-segmented version, and the segmentation tool and its version are stated.

  • Turkish Speech Translation

    State the alignment granularity (sentence / phrase / word level); word-level alignment needs stated rules for many-to-many cases.

  • Vietnamese Speech Translation

    The rule for handling pronoun information goes into the translator handbook, and the delivery notes describe the trade-off principle.

  • Filipino Speech Translation

    The rule for handling English content is uniform and written into the handbook; the delivery states the share of English content.

  • Persian Speech Translation

    Provide a spoken-to-written correspondence table for translators, and state the principle for handling spoken forms.

  • Tamil Speech Translation

    Translators must pass a Tamil spoken-comprehension test; a spoken-to-written correspondence table is provided.

  • Telugu Speech Translation

    Translator coverage must span the main dialect regions; state the sample size per dialect and each translator's dialect background.

  • Malayalam Speech Translation

    Whether the translation is done by the person who transcribed or by a second translator — Malayalam's fast speaking rate makes the two roles hard to combine, and the difference shows in the output.

  • Nepali Speech Translation

    Translators must pass a test; QC spot-checks for mistranslation caused by Hindi interference, with the spot-check records provided.

  • Sinhala Speech Translation

    Whether English proper nouns and technical terms are kept in English or transliterated into Sinhala script — both appear in the same recording and the choice changes the target text substantially.

  • Bangla Speech Translation

    State the source regional baseline and the target register separately, and deliver in batches by region.

  • Khmer Speech Translation

    State whether the source transcription uses orthography or phonemic notation; translators have to be able to see the corresponding actual pronunciation.

  • Arabic (Egyptian) Speech Translation

    Whether the target is English or Modern Standard Arabic, and whether the register of the source has to survive — rendering Egyptian colloquial as neutral formal text deletes the exact feature the data exists to capture.

  • Arabic (Gulf) Speech Translation

    Fix the orthography convention for the Gulf dialect source transcript, and require established published renderings for Quranic and religious passages rather than fresh translation.

  • Kannada Speech Translation

    Fix one source form for the whole project — raw colloquial transcript or normalized written standard — and if normalization is used, supply the expansion table translators must follow; habitual expansion shifts tense and politeness, and the error is invisible in the target text.

  • Punjabi Speech Translation

    State which script the source transcript uses, require translators to work from the audio rather than the transcript alone, and fix the target-language pair per side — tone is not recoverable from any transcript.

  • Burmese Speech Translation

    Write the rule for information the source does not state — tense, plurality, gender — into the translator handbook, and require translators to pass a colloquial comprehension test before the batch starts.

  • Amharic Speech Translation

    The handling of honorific forms goes into the translator handbook as an explicit principle, and the source transcription convention, including whether dropped sixth-order vowels are noted, is stated in the same document.

  • Ukrainian Speech Translation

    State the surzhyk rule for the source side — normalize first or translate as spoken — and attach the project glossary of recent vocabulary. All three components ship together.

  • Swahili Speech Translation

    State whether the source transcript is colloquial or normalized before translation begins, and give one uniform rule for English content in the source — all three components (audio, source transcript, target translation) ship together.

  • Kurdish Speech Translation

    The source branch, the transcript script, and the rule for state-language segments are all fixed before translation begins, and translators are branch-matched to the audio they handle.

  • Hausa Speech Translation

    One rule for Arabic formulaic phrases and one for English terms, applied uniformly across the batch, plus the target register (written-style or spoken-style Hausa) fixed in advance; the two read very differently.

  • Uzbek Speech Translation

    State the source transcript script and the rule for Russian-origin items — translated, kept, or mapped to the official coinage — in the translator handbook, with worked examples.

About Speech Translation →

In-the-Wild Speech

26 languages

  • Hindi In-the-Wild Speech

    Record the device (phone / headset / professional mic) and the collection environment for every item, or the data can be neither reproduced nor explained.

  • Arabic (MSA) In-the-Wild Speech

    Write out the collection task design, and state how artificial the resulting corpus is (whether speakers were asked to deliberately use the standard form).

  • Indonesian In-the-Wild Speech

    Annotate language composition and environmental conditions together, with neither normalized away.

  • Thai In-the-Wild Speech

    Annotate measured signal-to-noise ratio, and state which task types this data suits (tone modeling may not be among them).

  • Turkish In-the-Wild Speech

    Whether German and other loanwords common among Turkish speakers abroad are transcribed in Turkish orthography or in the source language — the two are not interchangeable for a language-ID-aware model.

  • Vietnamese In-the-Wild Speech

    Measure the intelligibility rate and report effective sample size from it (not raw recorded hours).

  • Filipino In-the-Wild Speech

    Language annotation and environment annotation are both done, measured separately, with neither standing in for the other.

  • Persian In-the-Wild Speech

    The speaker's regional variety must be logged per file — Tehrani, Esfahani, Mashhadi and Dari speakers are easy to pool and are not interchangeable for a dialect-aware model.

  • Tamil In-the-Wild Speech

    Annotators must pass a colloquial-comprehension test; the intelligibility rate is measured and reported honestly.

  • Telugu In-the-Wild Speech

    Annotate dialect region record by record; state the geographic range of collection and the sample size per dialect.

  • Malayalam In-the-Wild Speech

    Assess annotation hours from actual difficulty; the effective sample size is best derived by adjusting for the intelligibility rate.

  • Nepali In-the-Wild Speech

    Record device model, sampling rate, and duration for every item, and describe the quality distribution.

  • Sinhala In-the-Wild Speech

    Measure the actual share of each language component and present it honestly in the delivery notes.

  • Bangla In-the-Wild Speech

    Whether bystander speech is kept or cut, and if kept, whether the bystander's regional variety is labelled — in-the-wild Bangla recordings routinely capture a second speaker from a different variety.

  • Khmer In-the-Wild Speech

    The speaker's dialect origin has to be logged per file — urban Phnom Penh Khmer, rural varieties, and Khmer Krom or Surin speakers differ enough that pooling them unlabelled destroys the dialect signal.

  • Arabic (Egyptian) In-the-Wild Speech

    The transcript convention must state that Egyptian colloquial is written as spoken, with standard-language words appearing only if the speaker actually said them — and the plan has to name its regional coverage, because convenience sampling returns Cairo and nothing else.

  • Arabic (Gulf) In-the-Wild Speech

    Annotate device, environment, and the speaker's Arabic variety for every file, and report the measured share of Gulf versus non-Gulf speech in the delivery.

  • Kannada In-the-Wild Speech

    Annotate background speech by language where it is intelligible, with a stated exclusion threshold — urban Karnataka recordings routinely capture several languages in the background, and unlabelled background speech corrupts both language identification and diarization.

  • Punjabi In-the-Wild Speech

    Annotate the collection mode (self-recorded or field-transcribed by staff) and each speaker's script literacy — the two decide which parts of the population the dataset actually represents.

  • Burmese In-the-Wild Speech

    Log for every file whether Burmese is the speaker's first language and which region the recording was made in — without both, a mixed batch cannot be interpreted after the fact.

  • Amharic In-the-Wild Speech

    Log device model, sampling rate, and whether the recording passed through a messaging app before submission, for every item. A re-encoded file cannot be treated as equivalent to a direct recording.

  • Ukrainian In-the-Wild Speech

    Log device, environment, and location type for every item, and deliver a measured language-composition report covering how much of the batch is mixed rather than pure Ukrainian.

  • Swahili In-the-Wild Speech

    Every item logs the speaker's first language, the recording device, and the environment at collection time — first language is the one field that cannot be reconstructed later, and the delivery should state the intelligibility rate honestly.

  • Kurdish In-the-Wild Speech

    Device and environment are logged per item, and the delivery states whether state-language segments were kept and tagged, since the natural mixing is part of the recording rather than an error to clean up.

  • Hausa In-the-Wild Speech

    Record device, environment, and the speaker's first language for every item, and state whether second-language Hausa speakers are included by design or excluded; the two choices produce entirely different datasets under one category name.

  • Uzbek In-the-Wild Speech

    Log device, environment, and speaker origin per file, and measure the actual Russian-language share after collection — a preset ratio will not match what the recordings contain.

About In-the-Wild Speech →

Not seeing the combination you need?

These are the pairings where we have enough collection experience to describe the problem precisely. That is not the limit of what we can source. If your language or category is not listed, send the specification anyway — we will tell you whether we can source it well, and we would rather decline than deliver something that misses the bar.

All languages → All categories →

Request a dataset

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com