Global

Voice Anonymization

The short answer. Voice anonymization is a signal-processing problem with a legal standard attached. The techniques — pitch shifting, formant shifting, voice conversion — all change the signal, and all leave information a motivated attacker can use. The legal question is not whether the voice sounds different but the test in GDPR Recital 26: whether identification is reasonably likely, taking account of the means available, their cost and the time required. In practice most deployed anonymization is pseudonymization, and that distinction decides whether the GDPR applies at all.

The law

GDPR Recital 26

To determine whether a natural person is identifiable, account should be taken of all the means reasonably likely to be used, such as singling out, either by the controller or by another person, to identify the natural person directly or indirectly. To ascertain whether means are reasonably likely to be used to identify the natural person, account should be taken of all objective factors, such as the costs of and the amount of time required for identification, taking into consideration the available technology at the time of the processing and technological developments.

This is the standard an anonymized voice file has to meet, and it is deliberately not a technical threshold. Two consequences follow. The assessment moves with the state of the art, so a method that passed in 2020 may not pass now. And it accounts for means available to other people, not only to you, which includes the possibility that an enrollment sample of the speaker exists somewhere outside your control.

GDPR Article 32(1)(a)

the pseudonymisation and encryption of personal data

Read in context, Article 32 requires technical and organizational measures appropriate to the risk, and names pseudonymization as one of them. That placement is the point: pseudonymization is a security control, not an exit from the regulation. If a transformation does not meet the Recital 26 standard, what you have is a well-secured personal dataset with the full obligation set still attached.

Who it applies to

No statute requires anyone to anonymize a voice recording, and none prohibits it. The question is always what is being claimed. A vendor that describes a corpus as anonymized is making a legal claim, that the data falls outside the GDPR, and that claim is testable.

The Recital 26 test is applied by reference to the controller and to other people, which means the vendor's key file, a speaker manifest, a payment record, or an enrollment recording held by a third party all count as routes back. In a dataset context this produces a counterintuitive result: a single file can be genuinely hard to re-identify on its own and much easier to re-identify when delivered alongside five hundred hours from the same speaker, because the shared characteristics become the identifying signal.

The transcript layer is usually the weakest point, and it is the one buyers forget. Anonymizing the audio does nothing about a speaker who states their name, names their employer, describes their town, or uses a distinctive phrase that a search engine can match. If the corpus ships with transcripts, the anonymization claim covers less than it appears to.

What it costs to get wrong

There is no fine for using a weak anonymizer. The penalty is that nothing changed: if the output is still personal data, the lawful basis requirement, the data subject rights, the processor contract and the transfer rules all still apply, and a buyer who treated the corpus as outside the GDPR has been processing personal data without the instruments in place.

That lands on Articles 5, 6 or 9, which is the higher tier under Article 83(5): up to 20 million euros, or up to 4% of total worldwide annual turnover of the preceding financial year, whichever is higher. In an Illinois context the same failure is worse in structure, because BIPA damages are counted per person rather than per practice.

The contractual cost arrives sooner than either. A warranty that the data is anonymized, later shown to be wrong, sits at the center of the purchase. The buyer has paid for data it believed it could use freely, and the remedy is a claim against a vendor rather than a fix for the product the data already went into.

How to comply when you are buying data

The technical content below is what to ask about. The standard it has to meet is the legal one in Recital 26, and the two are not the same question.

  • Ask which method and at what settings. Pitch shifting moves the fundamental frequency. Formant shifting moves the vocal tract resonances. Voice conversion maps the recording onto a different speaker's characteristics using a model. They fail differently: speaker verification systems keyed to spectral envelope survive simple pitch shifts, while prosody, speaking rate and dialect survive formant changes.
  • Ask what evaluation was run: which attacker model, which metric, and whether the evaluation assumed the attacker holds an enrollment sample of the target speaker. An evaluation that assumes no enrollment sample measures an easier problem than the one that matters.
  • Check whether the transformation is applied per file or per speaker. A consistent transformation across one speaker's files can preserve a stable pseudo-identity that links the files to each other, which is a linking attack even without a name.
  • Read the transcript. If the corpus ships with text, assess the text separately, because content identifies speakers that audio transformation does not touch.
  • Treat "anonymized" as a claim to be tested rather than a property of the file, and write the vendor's answer into the contract as a warranty with a defined standard and a notification duty if the vendor's assessment changes.

How we handle consent and licensing →

Related compliance topics

Not legal advice

We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.

Sourcing data under Voice Anonymization?

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com