Compliance, consent and licensing
A dataset you cannot legally use is not a dataset. Compliance is not a document we attach at the end of a project — it is a constraint on how the project is designed, starting from speaker recruitment.
What every delivery includes
Regardless of language or category, a delivery from us comes with documentation that establishes where the data came from and what may be done with it.
- Signed speaker consent — States the intended use concretely, rather than a general release.
- Collection methodology — How speakers were recruited, screened and scheduled, so you can judge how biased the pool is.
- Annotation guideline — The document annotators worked from, so the conventions are reproducible.
- Recording conditions — Per file: environment, device, and signal-to-noise ratio where measured.
- Transfer mechanism — Where data crosses jurisdictions, addressed in the agreement rather than left implicit.
What we decline
Some categories carry consent exposure we are not set up to hold, and we turn those projects down rather than pass the risk downstream.
- Recorded telephone calls. The consent of the other party is typically absent, and retrofitting it is not possible.
- Medical and clinical data. The regulatory layer on top of consent is substantial, and we are not equipped for it.
- Children's voice data without guardian consent. Guardian consent is a hard requirement, and the permitted use is stated explicitly. We do not produce this data without it.
- Data with unclear provenance. If a buyer wants data whose origin cannot be documented, the answer is no.
Voice data is personal data
A voice recording can be linked to the person who made it, which makes it personal data under most data protection regimes. When it is used to identify a person, it is usually treated as biometric data as well, which is a stricter category with correspondingly heavier requirements.
The practical consequence is that consent for voice data has to describe the purpose specifically. A speaker who agreed to "be recorded" has not agreed to train a machine learning model or to have their voice used for identification.
What your legal team should ask for
- The consent form template, before production starts rather than at delivery.
- The lawful basis for processing, stated explicitly.
- Where the speakers are located, and where the data will be processed and stored.
- Whether the model trained on this data will be distributed, and whether the license covers that.
- How an erasure request would be handled once the dataset has been delivered.
We would rather answer these questions during scoping than during a dispute. Raise them early and the terms can be built to accommodate them.
What we cannot tell you
We are a sourcing company, not a law firm, and nothing on this page is legal advice. Whether a particular dataset is permissible in your jurisdiction depends on your use case, your location, and the location of the people in the recordings. Our role is to document the chain accurately so that your counsel can assess it.
Regulation by regulation
The rules that come up most often when buying speech and training data, each with the actual provision cited and what it means in practice.
Cross-border
-
AI Training Data Licensing
Training a model on someone else's work means making copies of it, and the reproduction right belongs to the rights holder. The EU has two text-and-data-mining exceptions, but one is limited to research organizations and the other can be switched off by the rights holder. The US has no training-specific exception at all. A license is how you replace that uncertainty with a document. The three things to check are which rights the grant actually names, whether it covers the compiled corpus as well as its contents, and whether it survives into the models you ship.
-
AI Training Data and Copyright
Two different acts get confused here. Ingesting a work to train a model is a reproduction, argued in the US as fair use and in the EU under a text-and-data-mining exception that the rights holder can switch off. Emitting something substantially similar to a specific training item is a separate act, and it is much harder to defend. No US appellate court has settled whether training is fair use. Treat any vendor who tells you it is settled as describing a litigation position rather than a clearance, and check which of the two acts your product actually performs.
-
Data Provenance
A provenance record is worth having only if it survives to the file. "Collected from the internet" is not provenance, because it answers where the data came from without answering what was done to it, under what authority, and who agreed. A usable record carries the source, the collection date and method, the rights basis, the consent reference, the transformations applied, and the retention terms. Two provisions make this concrete rather than philosophical: the training-content summary required of general-purpose model providers, and the record of processing activities required of controllers.
-
Data Licensing Agreement
Five clauses decide whether a data deal works: the scope of use, sublicensing and model-distribution rights, term and termination, indemnity, and what happens to models already trained when the license ends. The last one is the clause most often missing and the most expensive to be without, because a model cannot be untrained. Where the corpus contains personal data, the agreement also has to do the work of a data processing contract, because GDPR does not let a person sign away their rights the way a copyright owner can license a work.
-
AI Training Data Laws
There is no single global rule on training data, and any page that offers one is describing a jurisdiction rather than the law. Four regimes interact: the EU AI Act, which imposes documentation and transparency duties on model providers; the GDPR, which governs the people in the data; US copyright law, where the training question is litigated under fair use rather than settled by statute; and a growing set of US state laws that mostly require disclosure rather than permission. Your obligation depends on where you operate, where you sell, and where the speakers were.
-
AI Training Data Governance
A governance process for training data is four controls, not a policy document: a lineage record that ties every delivered file to a named source and a consent artifact; an approval gate that data has to pass before it enters a training run; a retention rule that says when the data is destroyed; and a procedure for acting on an erasure request after the model has already trained. GDPR Article 5(1)(e) supplies the retention rule. EU AI Act Article 10 supplies the quality and provenance standard. Everything else is implementation.
-
Vendor Due Diligence Checklist
Due diligence on a data vendor is the process of testing six things before you pay: the consent artifacts, where the speakers were recorded, the chain of custody from recording to delivery, the sub-processor list, the security posture, and what happens when a speaker asks for their data to be deleted. GDPR Article 28(1) states the underlying rule, that a controller may use only processors providing sufficient guarantees. Article 5(2) puts the burden of demonstrating compliance on the controller, which is you.
-
Voice Anonymization
Voice anonymization is a signal-processing problem with a legal standard attached. The techniques — pitch shifting, formant shifting, voice conversion — all change the signal, and all leave information a motivated attacker can use. The legal question is not whether the voice sounds different but the test in GDPR Recital 26: whether identification is reasonably likely, taking account of the means available, their cost and the time required. In practice most deployed anonymization is pseudonymization, and that distinction decides whether the GDPR applies at all.
-
Ethical Data Collection
Ethical data collection is the set of practices that go beyond the legal minimum: paying speakers at a rate that reflects the session, telling them what the model is for before they agree, not recording people in public without notice, and documenting the terms that were actually explained rather than the terms on the form. EU AI Act Article 10 gives the buyer a reason to care, because the provider has to document the origin of the data and the original purpose of collection. GDPR Article 13 requires the transparency in the first place.
-
Consent for AI Training
Consent taken for one purpose does not automatically cover training a model. GDPR Article 5(1)(b) requires data to be collected for specified, explicit and legitimate purposes and not further processed incompatibly, and Article 4(11) requires consent to be specific and informed. Where the original basis was consent, the practical test is whether the speaker, at the moment of recording, would have expected their voice to end up in a model that may be distributed. If the form did not say so, the answer is usually no.
-
Children's Voice Datasets
Two regimes apply and they are not the same. GDPR Article 8 sets conditions for a child's consent in relation to information society services, with the age of digital consent set by each member state between 13 and 16. In the US, COPPA and its rule at 16 C.F.R. Part 312 govern collection from children under 13, and the rule treats an audio file containing a child's voice as personal information by definition. Guardian consent is a hard requirement, and the permitted use has to be stated explicitly.
-
Synthetic Data
Synthetic data solves three real problems: there is no consent chain to document, no speaker to re-identify, and the balance of the corpus can be set deliberately. It does not solve everything. A synthetic corpus inherits the biases and the recording conditions of the real data it was modeled on, and if it was generated from personal data the GDPR question does not disappear. A synthetic record is outside the regulation only where re-identification is not reasonably likely.
European Union
-
Voice Data under GDPR
A voice recording of an identifiable person is personal data, and holding it is processing. That means you need a lawful basis, and for AI training the realistic basis is consent or legitimate interests. You also need the speaker to have been told the specific purposes, because Article 13 puts that duty on whoever collected the data, and a buyer inherits the gap if it was never done. Where the voice is used to identify people, Article 9 raises the bar further. Check the basis, the notice and the transfer mechanism before you take delivery.
-
GDPR Consent for Voice Recording
Valid consent under GDPR is four things at once: freely given, specific, informed and unambiguous. A blanket release that covers "recording" does not cover training a model on the recording, because specificity is measured purpose by purpose. Consent also has to be separable from other terms, and as easy to withdraw as it was to give. Most voice dataset consent forms fail on specificity rather than on signature. Read the actual form text before you accept a dataset, not a summary of it.
-
Is Voice Personal Data?
Yes. A recording of a speaker who can be identified is personal data, because it relates to an identifiable natural person, and the GDPR definition expressly includes factors specific to the physical and physiological identity of that person. It stays personal data if you never run speaker recognition on it, and it stays personal data if you replace names with speaker codes. Simply storing the audio is itself processing. The practical test is whether the speaker can be identified by means reasonably likely to be used, not whether you intend to identify them.
-
Is Voice Biometric Data?
It depends on what you do with it, not on what it is. Under GDPR, voice is biometric data only when it is processed through specific technical means for the purpose of uniquely identifying a person. Training an ASR model on voice is generally not that, because the purpose is transcription. Building a speaker identification or verification system is, because unique identification is the whole point. That distinction decides whether Article 9 applies, and it is usually the buyer's intended use rather than the supplier's collection method that settles it.
-
EU AI Act Transparency Requirements
Two different transparency regimes run side by side, and they are easy to confuse. Article 50 imposes disclosure duties on the providers and deployers of certain AI systems — telling people they are talking to a machine, marking synthetic output, and disclosing emotion recognition. Article 53 imposes documentation and disclosure duties on providers of general-purpose models, including a public summary of training content and a copyright policy. The dates are phased: general-purpose obligations applied from 2 August 2025 and the Article 50 duties from 2 August 2026.
-
EU AI Act Article 50
Article 50 is the transparency article, and it contains four separate duties. Providers must tell people they are interacting with an AI system, and must mark synthetic audio, image, video and text output as machine-readable and detectable. Deployers must inform people exposed to emotion recognition or biometric categorisation, and must disclose deepfakes and AI-generated text on matters of public interest. Each duty has its own exceptions, and the disclosure has to come at the latest at first interaction or exposure. It applied from 2 August 2026.
-
Biometric Data under GDPR
Voice is not automatically special-category data under the GDPR. Article 4(14) defines biometric data as personal data resulting from specific technical processing of physical, physiological or behavioural characteristics that allow or confirm unique identification, and Article 9(1) prohibits processing of biometric data for the purpose of uniquely identifying a natural person. The test is the purpose, not the file format. The same recording is ordinary personal data in a transcription corpus and special-category data in a speaker verification corpus.
-
Data Protection Impact Assessment
A DPIA is required before processing that is likely to result in a high risk to the rights and freedoms of natural persons, under GDPR Article 35(1). Article 35(3) lists the processing that always requires one, and the item that catches voice data at scale is large-scale processing of the special categories in Article 9(1). The assessment has to contain the four elements in Article 35(7). If the residual risk stays high after mitigation, Article 36 requires consultation with the supervisory authority before processing begins.
-
Data Processing Agreement
If you buy a dataset containing personal data and you decide what it is used for, you are a controller and the vendor is a processor, and the arrangement has to be governed by a contract meeting GDPR Article 28(3) rather than by a license. Article 28(3) sets out what the contract must stipulate. Sub-processors need authorization under Article 28(2) and the same obligations have to flow down under Article 28(4). If the data moves outside the EU, Chapter V applies, usually through the standard contractual clauses under Article 46.
-
Anonymisation vs Pseudonymisation
Pseudonymised data is still personal data. GDPR Recital 26 says so directly, and Article 4(5) defines pseudonymisation as processing that prevents attribution without additional information. Anonymised data is outside the regulation entirely. For a dataset purchase that is the difference between a full obligation set — lawful basis, data subject rights, a processor contract, transfer rules — and none of them. A vendor's claim that a corpus is anonymised is a legal claim, tested against Recital 26 rather than against the label on the file.
United States
-
Biometric Privacy Laws
There is no US federal biometric privacy statute. The law is state by state, and only a handful of states have one. Illinois, Texas and Washington are the long-standing three, Colorado joined with an amendment effective July 1, 2025, and more states are adding biometric provisions to general privacy acts rather than writing standalone statutes. Which ones matter depends on two things: whether the statute names voiceprints, and whether it is triggered by collection or only by identification. Illinois and Texas name voiceprints. Washington expressly excludes audio recordings.
US state law
-
Two-Party Consent States
Federal law allows a participant to record a conversation, and states may be stricter. Around a dozen require the consent of all parties, and the commonly cited list is California, Connecticut, Florida, Illinois, Maryland, Massachusetts, Michigan, Montana, New Hampshire, Oregon, Pennsylvania and Washington, with Nevada often added. Several of those have carve-outs by communication type, and Michigan's participant rule is genuinely contested. Treat the list as the start of an inquiry. For a dataset, the safe position is all-party written consent wherever the speaker was, because the strictest state in the chain sets the standard.
-
Voice Recording Consent by State
Wiretap consent and dataset consent are different questions, and clearing the first does not clear the second. A state's recording law decides whether you were allowed to make the recording. What a data buyer needs is the speaker's agreement to the training and distribution uses, which comes from consent law and right-of-publicity law rather than from the wiretap statutes. The state that matters is where the speaker was, not where your company sits. Three states — California, New York and Tennessee — name voice expressly in their publicity statutes.
-
California AI Training Data Transparency Act
AB 2013 (2024) added Title 15.2 to the California Civil Code. Section 3111 requires the developer of a generative AI system or service made available to Californians to post documentation about the data used to train it, on or before January 1, 2026 and before each subsequent release or substantial modification. It is a public summary of the dataset, not a filing of the dataset, and the duty falls on the developer rather than on the data vendor who supplied the recordings.
-
Illinois BIPA
BIPA is an Illinois statute, 740 ILCS 14/, and a voiceprint is one of the biometric identifiers it lists. A private entity that collects a voiceprint from an Illinois resident must give written notice, state the specific purpose and the length of term, and obtain a written release before collection, under §14/15(b). It must publish a retention and destruction schedule under §14/15(a) and may not disclose the data without consent under §14/15(d). The 2024 amendment allows an electronic signature to serve as the written release.
Related
-
Data compliance
What the term covers, and why it is a design constraint rather than paperwork.
-
GDPR
The EU regulation, and the two parts that matter most for voice data.
-
Consent
Why a general release is not sufficient for AI training.
-
AI data licensing
What the license actually permits, and the clauses that decide whether a deal works.
Submit a sourcing request
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.