EU
Is Voice Personal Data?
The short answer. Yes. A recording of a speaker who can be identified is personal data, because it relates to an identifiable natural person, and the GDPR definition expressly includes factors specific to the physical and physiological identity of that person. It stays personal data if you never run speaker recognition on it, and it stays personal data if you replace names with speaker codes. Simply storing the audio is itself processing. The practical test is whether the speaker can be identified by means reasonably likely to be used, not whether you intend to identify them.
The law
GDPR Article 4(1)
"personal data" means any information relating to an identified or identifiable natural person ("data subject"); an identifiable natural person is one who can be identified, directly or indirectly, in particular by reference to an identifier such as a name, an identification number, location data, an online identifier or to one or more factors specific to the physical, physiological, genetic, mental, economic, cultural or social identity of that natural person.
Voice sits inside the definition twice over. The recording relates to a person, and the voice itself is a factor specific to their physical and physiological identity. Nothing in the definition requires the data to be used for identification, so the answer does not change based on what you plan to build.
GDPR Article 4(2)
"processing" means any operation or set of operations which is performed on personal data or on sets of personal data, whether or not by automated means, such as collection, recording, organisation, structuring, storage, adaptation or alteration, retrieval, consultation, use, disclosure by transmission, dissemination or otherwise making available, alignment or combination, restriction, erasure or destruction.
Storage is on the list. So is structuring, which is what building a manifest or an index of a corpus amounts to. A buyer who receives a voice dataset and holds it without doing anything else is already processing personal data, and therefore already needs a lawful basis for that holding.
GDPR Recital 26
To determine whether a natural person is identifiable, account should be taken of all the means reasonably likely to be used, such as singling out, either by the controller or by another person, to identify the natural person directly or indirectly. ... The principles of data protection should therefore not apply to anonymous information, namely information which does not relate to an identified or identifiable natural person or to personal data rendered anonymous in such a manner that the data subject is not or no longer identifiable.
The test is "reasonably likely means", which is a factual question about the data as it stands, not about your intentions. Replacing a speaker's name with a code produces pseudonymised data, and Recital 26 says explicitly that pseudonymised data which can be attributed to a person using additional information is still personal data. A corpus of raw audio with speaker codes is pseudonymised, not anonymous.
Who it applies to
The question is asked most often by teams holding a corpus they did not collect, which is the situation of every data buyer. You are not the person who decided to record these speakers, but once the files are in your storage you are processing personal data relating to them, and the analysis of your own basis is yours to make.
The misconception worth naming is that anonymisation happens automatically when you remove obvious identifiers. It does not. Removing names, dates of birth and contact details from a manifest produces pseudonymisation, and Recital 26 treats that as personal data. Genuine anonymisation requires that the speaker is no longer identifiable by anyone, using means reasonably likely to be used, which for raw voice is difficult to establish because the voice is the identifier.
The test has a second half that is easy to skip: account is taken of the means reasonably likely to be used, by the controller or by another person. A dataset can therefore be personal data because of what someone else could do with it. Removing the names does not settle the question while the voices remain.
One practical consequence is easy to miss. Because a voice recording is personal data, data subject rights attach to it, and access and erasure requests have to be operable against a distributed corpus. A dataset that cannot answer an erasure request has a defect in its design, not in its paperwork.
- Studio recordings of consenting adult speakers: personal data, straightforward to document.
- Pseudonymised corpora with speaker codes: still personal data, because the code maps back to a person.
- Genuinely anonymous audio where no one can single out a speaker: outside GDPR, and hard to demonstrate for raw voice.
What it costs to get wrong
Misclassifying personal data as anonymous removes the obligations without removing the exposure. There is no lawful basis, no Article 13 notice and no erasure mechanism, which puts the processing in the Article 83(5) tier at up to 20 million euros or 4% of total worldwide annual turnover, whichever is higher.
A separate provision applies if you keep records. Article 30 requires a record of processing activities, and failing to maintain one is in the lower tier under Article 83(4)(a) at up to 10 million euros or 2%. That is the smaller number, and it is the one that surfaces first in an audit, because the record is the document that proves you thought about the basis at all.
The commercial consequence is often the decisive one. Enterprise buyers in the EU will not accept a voice corpus whose supplier cannot say whether the data is personal, and the answer "we removed the names" is treated as a red flag rather than a reassurance. Deals are lost at that point, not at the point of a fine.
For a distributed corpus there is also an exposure that compounds: every copy already delivered is a separate processing operation, so a defect in the original classification propagates to each recipient rather than staying with the seller.
How to comply when you are buying data
The useful question is not "is this personal data in the abstract" but "what exactly is in the files, and who could be singled out from them". Answering it precisely is a documentation exercise, and the documentation is what a buyer will be asked to show.
Two answers are common and both are wrong. "It is just audio, we removed the names" describes pseudonymization, which the regulation treats as personal data. "It is public data, so it is not personal" confuses availability with identifiability, which are unrelated questions.
Where a project involves EU speakers, we document the classification and the reasoning rather than asserting a conclusion, and the consent record travels with the dataset. We are a sourcing company and not a law firm, so the classification decision for your specific use remains one for your counsel; what we can do is make the facts complete enough for that decision to be made on evidence.
- Ask what identifiers travel with the files. Filenames, metadata fields, session logs and the manifest are all part of the data, and a stripped audio file with a name in its filename is not anonymous.
- Ask whether the supplier claims anonymisation or pseudonymisation, and on what basis. A supplier who uses the two words interchangeably has not done the assessment.
- Ask for the record of processing activities if you are acting as a controller. If it does not exist, the cheapest first step is to create one for the data you now hold.
- Confirm the erasure path. A speaker identifier on every file is what makes an erasure request actionable; without it the request cannot be serviced at all.
- Check the transfer basis if the corpus crosses borders, because personal data crossing a border is a restricted transfer and anonymity is the only thing that would remove it.
Related compliance topics
-
Is Voice Biometric Data?
is voice biometric data
-
Biometric Privacy Laws
biometric privacy laws
-
Two-Party Consent States
two party consent states
-
Voice Recording Consent by State
voice recording consent by state
-
EU AI Act Transparency Requirements
eu ai act transparency requirements
-
EU AI Act Article 50
eu ai act article 50
Not legal advice
We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.
Sourcing data under Is Voice Personal Data??
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.