EU

Anonymisation vs Pseudonymisation

The short answer. Pseudonymised data is still personal data. GDPR Recital 26 says so directly, and Article 4(5) defines pseudonymisation as processing that prevents attribution without additional information. Anonymised data is outside the regulation entirely. For a dataset purchase that is the difference between a full obligation set — lawful basis, data subject rights, a processor contract, transfer rules — and none of them. A vendor's claim that a corpus is anonymised is a legal claim, tested against Recital 26 rather than against the label on the file.

The law

GDPR Recital 26

Personal data which have undergone pseudonymisation, which could be attributed to a natural person by the use of additional information should be considered to be information on an identifiable natural person.

The additional information is the crux. A voice file with the speaker's name replaced by a code is pseudonymised if the mapping exists anywhere, including in the vendor's own records, and the regulation says plainly that this remains personal data. If the mapping is destroyed and no other route to identification is reasonably likely, the analysis changes — but the burden sits on the party making the claim.

GDPR Article 4(5)

'pseudonymisation' means the processing of personal data in such a manner that the personal data can no longer be attributed to a specific data subject without the use of additional information, provided that such additional information is kept separately and is subject to technical and organisational measures to ensure that the personal data are not attributed to an identified or identifiable natural person.

The definition carries two conditions: the additional information is kept separately, and it is subject to technical and organizational measures preventing attribution. A spreadsheet of speaker codes and names sitting in the same vendor's storage fails the second condition, which means the delivery is not even pseudonymised to the standard the definition describes.

GDPR Article 4(1)

'personal data' means any information relating to an identified or identifiable natural person ('data subject'); an identifiable natural person is one who can be identified, directly or indirectly, in particular by reference to an identifier such as a name, an identification number, location data, an online identifier or to one or more factors specific to the physical, physiological, genetic, mental, economic, cultural or social identity of that natural person.

This is where the whole question sits. Anonymised data is not personal data, so Article 4(1) does not reach it and nothing else in the regulation does either. That is why the distinction is commercial and not academic: an anonymised corpus can be sold across borders, kept indefinitely and used without a lawful basis, and a pseudonymised one cannot do any of those things.

Who it applies to

The distinction is a factual determination about a specific dataset in specific hands, and it is not settled by contract. A clause stating that the parties agree the data is anonymised does not make it so; supervisory authorities and courts look at the data and the means available, not at the recital of the agreement.

Nor is the answer necessarily the same for both parties. The test turns on the means reasonably likely to be used by the controller or by another person, so the same file can sit at the edge of personal data for a buyer that holds nothing else, and be clearly personal data for the vendor that holds the consent forms, the recording schedule and the payment records. A buyer that assumes its own position settles the question for the vendor is making an assumption the regulation does not support.

For procurement, the practical scope question is what the vendor retains. The mapping key, the speaker manifest, the raw session files, the recruitment records and the payment ledger are each a route back to the speaker. If any of them still exists, the delivery is pseudonymised with respect to the vendor, and the anonymisation claim is about a narrower thing than the buyer probably assumes.

What it costs to get wrong

The consequence is not a separate offense. It is that the obligations you believed did not apply do apply. If the corpus is pseudonymised, then a lawful basis was needed for the training use, access and erasure requests have to be answerable, an Article 28 contract with the vendor is required, and any transfer out of the EU needed a mechanism.

Missing those puts the processing under Articles 5, 6 or 9, which is the higher tier under Article 83(5): up to 20 million euros, or up to 4% of the total worldwide annual turnover of the preceding financial year, whichever is higher.

The commercial penalty usually arrives first. An enterprise buyer that discovers a delivered corpus is pseudonymised rather than anonymised has to unwind the product it went into, because the lawful basis, the retention schedule and the disclosure record for that product were built on the wrong premise. That cost lands on the purchase contract, and the indemnity is where it is argued.

How to comply when you are buying data

The work is to convert a label into a documented position with a named standard.

  • Ask the vendor to state in writing whether the corpus is anonymised or pseudonymised, and to describe the transformation in enough detail to be checked.
  • Ask what the vendor retains: the mapping key, the speaker manifest, the raw recordings, the consent forms, the recruitment records. If any of them exists, the delivery is pseudonymised with respect to the vendor.
  • Test the claim on a sample. Take two files from the same speaker delivered as separate records and see whether the vendor's own metadata links them. That is the linking test, and it takes minutes.
  • Look at the text layer. Even with the audio transformed, a transcript can identify a speaker through content alone, which means the anonymisation claim has to be assessed per layer rather than per delivery.
  • If the claim is anonymisation, take it as a warranty with a defined standard and a notification duty if the vendor's assessment changes. Technology moves, and so does the Recital 26 answer.
  • If the answer is pseudonymisation, price the obligations in from the start: the DPA, the impact assessment, the retention rule, the erasure procedure and the transfer mechanism are all part of the cost of the dataset.

How we handle consent and licensing →

Related compliance topics

Not legal advice

We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.

Sourcing data under Anonymisation vs Pseudonymisation?

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com