What is Anonymization?

Anonymization — removing or altering identifying information so a person can no longer be identified, in principle irreversibly, which is what separates it from pseudonymization.

Anonymization means processing personal data so that it can no longer be linked to an identifiable person. The bar is high: deleting names is not enough. Voice, face, gait, writing style, and combinations of apparently harmless details can all re-identify someone, and a dataset that looks anonymous can often be de-anonymized by joining it against another dataset.

This is why regulators treat anonymization as a result rather than a technique. Under the GDPR, data that has genuinely been anonymized falls outside the regulation entirely, because it is no longer personal data. Data that has merely been pseudonymized — identifiers replaced with codes, but a key retained somewhere — remains personal data and stays fully in scope.

For voice the difficulty is structural. A recording carries identity in a way a text file does not, so removing identity without destroying the signal a model needs is genuinely hard. Speech anonymization usually works by altering the characteristics that carry identity while preserving the linguistic content, and the result is a tradeoff: the more identity is removed, the more naturalness and speaker variety the buyer loses.

The practical consequence for a supplier is that anonymization is not a substitute for consent. If the data was collected without a rights chain, anonymizing it afterwards does not create one — and if the anonymization turns out to be reversible, the data was personal data all along.

Related terms

  • PII

    Personally Identifiable Information — anything that can identify a specific individual, directly or indirectly.

  • Consent

    The speaker knows how their recording will be used, and has signed to say so.

  • Biometric Data

    Biometric Data — data derived from a person's physical or behavioral characteristics, such as voice or face, which most privacy laws treat as a special and more tightly restricted category.

Keep reading

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com