Global

Children's Voice Datasets

The short answer. Two regimes apply and they are not the same. GDPR Article 8 sets conditions for a child's consent in relation to information society services, with the age of digital consent set by each member state between 13 and 16. In the US, COPPA and its rule at 16 C.F.R. Part 312 govern collection from children under 13, and the rule treats an audio file containing a child's voice as personal information by definition. Guardian consent is a hard requirement, and the permitted use has to be stated explicitly.

The law

GDPR Article 8(1)

Where point (a) of Article 6(1) applies, in relation to the offer of information society services directly to a child, the processing of the personal data of a child shall be lawful where the child is at least 16 years old. Where the child is below the age of 16 years, such processing shall be lawful only if and to the extent that consent is given or authorised by the holder of parental responsibility over the child. Member States may provide by law for a lower age for those purposes provided that such lower age is not below 13 years.

Read the scope carefully. Article 8 applies to information society services offered directly to a child, so a recording session is not automatically inside it. Outside that context, consent for a child is given by the holder of parental responsibility as the child's legal representative, under Article 6(1)(a) read with national law on capacity. Either way the operational requirement is the same: the guardian consents, and the consent has to be as specific as an adult's.

15 U.S.C. §6501 et seq. — 16 C.F.R. §312.5(a)(1)

An operator is required to obtain verifiable parental consent before any collection, use, or disclosure of personal information from children, including consent to any material change in the collection, use, or disclosure practices to which the parent has previously consented.

Verifiable is a standard, not a word. The FTC's approved methods include a signed consent form returned by mail or electronic scan, a credit card or government ID check, and a call to a trained operator. A checkbox on a web page is not verifiable parental consent. Note also that a material change to the use requires fresh consent, which means expanding a corpus from transcription to speech synthesis is a new consent event.

16 C.F.R. §312.2 (definition of 'personal information')

A photograph, video, or audio file where such file contains a child's image or voice.

This is the provision that puts children's voice corpora squarely inside the COPPA rule: an audio file containing a child's voice is personal information by definition, without any analysis of whether it could identify anyone. The same section separately lists voiceprints as a biometric identifier. The rule also now limits retention — 16 C.F.R. §312.10 prohibits indefinite retention and requires a written retention policy — so a children's corpus needs a destruction plan from the day it is collected.

EU AI Act Article 5(1)(b)

the exploitation of any of the vulnerabilities of a natural person or a specific group of persons due to their age, disability or a specific social or economic situation…

A prohibition rather than a governance requirement. The full provision bars systems that exploit vulnerabilities due to age where the objective or the effect is to materially distort behaviour in a way that causes or is likely to cause significant harm. It binds providers and deployers directly, and it is not conditioned on a risk classification. A voice model built to imitate a specific child, or a system that targets children because of how they speak, is where that line is drawn, and the assessment is about the effect on the child rather than the developer's intent.

Who it applies to

COPPA binds operators of websites or online services directed to children, and operators that have actual knowledge they are collecting personal information from children. A recording application, a game, or a platform that captures children's voices is covered, and so is a vendor that collects through any of those channels. A purely offline recording session is not the target of the COPPA rule, which is a gap that state law and the GDPR fill rather than a permission.

The GDPR applies whenever the child is in the EU, whatever the collection method and wherever the buyer is. Several US states also require parental consent for processing children's data, and their thresholds differ from the federal one, so a corpus collected from thirteen-year-olds may be lawful under COPPA and outside a state rule that covers under-sixteens.

The AI Act provision is different in kind from the rest. It is a prohibition, it binds providers and deployers directly, and it does not depend on a risk classification or on a data governance record. For a children's corpus, that means the intended use has to be assessed against the prohibition before the collection is scoped, not after the model is built.

What it costs to get wrong

The FTC can seek civil penalties for violations of the COPPA rule. The amount is adjusted for inflation each year and has been above 50,000 dollars per violation in recent years, and state attorneys general can bring the same claims. Because the rule counts violations per collection, a corpus of several thousand children is a large number before any multiplier is applied.

Under the GDPR, a violation involving Article 8 sits in the Article 83(4) group at up to 10 million euros or 2% of total worldwide annual turnover, whichever is higher. A missing lawful basis for the underlying processing is the higher tier under Article 83(5), at up to 20 million euros or 4%.

A violation of the AI Act prohibition in Article 5 is the top tier under Article 99(3): up to 35 million euros, or up to 7% of total worldwide annual turnover, whichever is higher.

Beyond the penalties, children's corpora carry a specific commercial risk. A buyer that discovers the guardian consent was obtained through an unverifiable method has to remove the data from the model, and that is a retraining event with the full cost attached.

How to comply when you are buying data

Guardian consent is a precondition rather than a formality, and the permitted use has to be explicit enough that a guardian can give an informed yes or no.

  • Use a verifiable consent method: a signed form returned by mail or scan, a credit card or government ID check, or a recorded call with a trained operator. A checkbox is not one of them.
  • State the permitted use explicitly in the form: model training, distribution of the model, sublicensing, and whether the child's voice may be used to synthesize speech. If the answer to any of those is no, the form has to say so, because a guardian cannot consent to an unnamed use.
  • Name the retention period and the destruction date. The COPPA rule prohibits indefinite retention, and the GDPR storage limitation principle applies to the same data, so the two obligations point the same way.
  • Give the guardian a withdrawal route that works in practice, and record the withdrawals with dates. A withdrawal mechanism that requires a phone call during business hours is not a mechanism.
  • Ask the vendor for the consent method, not only the consent form. The method is what the verifiability standard tests, and it is the part that cannot be reconstructed later.
  • Do not accept a corpus where the speakers' ages are unverified. If the vendor cannot show how ages were checked, the corpus is a children's corpus you cannot document, whatever the file metadata says.

How we handle consent and licensing →

Related compliance topics

Not legal advice

We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.

Sourcing data under Children's Voice Datasets?

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com