Script splits: when one spoken language needs two datasets

Hindi and Urdu are mutually intelligible when spoken and mutually unreadable when written. That single fact reshapes how the data has to be built.

A Punjabi speaker in Lahore and a Punjabi speaker in Amritsar can hold a conversation without difficulty. Ask them to write it down and they will produce text in two scripts that share almost no characters.

This is a script split: one spoken language, two writing systems. It appears in Punjabi, Hindi and Urdu, Serbian and Croatian, Kazakh, Mongolian, Azerbaijani, and several others — and it changes what "a dataset for this language" means.

Why it is a data problem and not just a writing problem

Speech data is a pairing of audio with text. If the same audio can legitimately be paired with two different texts, then the dataset is not one dataset.

Train a model on Hindi-script text and deploy it where users type in Urdu script, and the acoustic model may be fine while the language model fails completely. The failure looks like a model problem. It is a specification problem.

The three ways to handle it

There is no universally correct answer, but there are three coherent options and the specification needs to pick one.

  • One script only — collect and annotate in the script your deployment target uses. Simplest, and correct when you know where the model will run.
  • Parallel scripts — annotate the same audio in both scripts, producing two aligned text corpora from one recording. Roughly doubles annotation cost, and gives you a genuinely bilingual asset.
  • Transliteration layer — annotate in one script and generate the other by rule. Cheapest, but transliteration between split scripts is not fully reversible, and the errors are systematic rather than random.

The related trap: script is not the same as language

Serbian and Croatian are the same language by most linguistic measures, written in Cyrillic and Latin respectively. Hindi and Urdu differ more in vocabulary than in grammar, but their scripts diverge completely.

This means a dataset labeled by script and a dataset labeled by spoken variety will not match. Buyers who specify "Serbian" and receive a dataset that is 40% Latin-script Ekavian will find that the label did not mean what they assumed.

How to specify it

State the script explicitly in the specification, alongside the language name. Then state whether transliteration is permitted, and if so, whether it must be human-verified.

This is a two-line addition to a specification that routinely prevents a full re-annotation.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com