Global

Synthetic Data

The short answer. Synthetic data solves three real problems: there is no consent chain to document, no speaker to re-identify, and the balance of the corpus can be set deliberately. It does not solve everything. A synthetic corpus inherits the biases and the recording conditions of the real data it was modeled on, and if it was generated from personal data the GDPR question does not disappear. A synthetic record is outside the regulation only where re-identification is not reasonably likely.

The law

GDPR Recital 26

To determine whether a natural person is identifiable, account should be taken of all the means reasonably likely to be used, such as singling out, either by the controller or by another person, to identify the natural person directly or indirectly.

The standard is applied to the output, not to the method. A synthetic corpus generated from five hundred real speakers is outside the GDPR if a record cannot reasonably be traced to one of them, and inside it if it can. Generation models can memorize, and a synthetic record that reproduces a distinctive phrase, a name, or an idiosyncratic delivery is a route back to the source even when no real file was copied.

EU AI Act Article 10(3)

Training, validation and testing data sets shall be relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose.

The standard a synthetic corpus has to meet if it substitutes for real data in a high-risk system. Representativeness is the difficult one: a generator trained on a narrow real corpus produces a large synthetic corpus with the same narrow distribution, and the size of the output disguises the narrowness of the input. The same provision requires appropriate statistical properties with regard to the people the system is intended to be used on, which is a question about coverage rather than volume.

GDPR Article 5(1)(d)

accurate and, where necessary, kept up to date; every reasonable step must be taken to ensure that personal data that are inaccurate, having regard to the purposes for which they are processed, are erased or rectified without delay ('accuracy').

The accuracy principle is the one buyers forget. If a synthetic record is traceable to a real person, it is arguably personal data about that person, and if it contains an utterance they never made — which is what synthetic generation produces — the accuracy principle applies to it. That objection is separate from identifiability and survives a successful anonymization argument.

Who it applies to

Synthetic data is a technique rather than a jurisdiction, so the scope question is what it is used for. If a synthetic corpus is generated from recordings of EU speakers, the GDPR question is asked about the outputs. If the generator is a model trained on personal data, that training is itself a processing operation with its own lawful basis, which is easy to overlook because the output looks like nothing.

Where the synthetic data feeds a high-risk AI system, Article 10 applies to it in the same terms as real data. The regulation does not create a lighter regime for generated content, and a provider cannot meet the representativeness requirement by pointing at the volume of synthetic records.

The commercial scope note is that synthetic data is usually bought as a supplement rather than a replacement. It is strongest where real data is scarce, legally blocked, or unbalanced in a way that cannot be fixed by recording more of the same. It is weakest where the acoustic conditions of the real world are the thing the model has to learn, because a generator reproduces the conditions it was trained on and smooths over the ones it never saw.

What it costs to get wrong

No provision penalizes synthetic data as such. Three risks sit around it.

The first is mislabeling. If synthetic data is delivered as real, or real as synthetic, the disclosures that turn on the source of the data are wrong. California's Civil Code §3111(a)(12) asks specifically whether synthetic data generation was used, and EU AI Act Article 53(1)(d) requires a public summary of the content used to train a general-purpose model. A wrong answer in either is a documented misstatement rather than a technicality.

The second is re-identification. If the outputs turn out to be personal data after all, the processing sits under Articles 5 and 6, which is the higher tier under Article 83(5): up to 20 million euros, or up to 4% of total worldwide annual turnover of the preceding financial year, whichever is higher.

The third cost shows up in evaluation rather than in a fine. A model trained on synthetic data that inherited a narrow distribution fails on real traffic, and the failure is usually discovered after the training budget is spent, which makes it the most expensive of the three.

How to comply when you are buying data

The questions below separate a synthetic corpus that solves a problem from one that moves it.

  • Ask what the generator was trained on and whether the seed data included personal data. The answer decides whether the GDPR question is closed or still open, and it is the first question rather than the last.
  • Ask for the re-identification evaluation: which attack was run, which metric was used, and whether the evaluation assumed the attacker holds the seed corpus. An evaluation against an attacker with no seed data measures an easier problem.
  • Ask for the balance report. A synthetic corpus should be able to state its speaker, accent, age and recording-condition distribution. If it cannot, it is a size figure with a method attached.
  • Ask whether the generator memorizes. Generative models can reproduce training examples, and a memorized record is a real record with a real speaker behind it, whatever the pipeline calls it.
  • Test the synthetic corpus against real evaluation data before committing to a full training run. Synthetic data is easiest to validate against the distribution it came from and hardest to validate against the one it will meet in production.
  • Keep the real and synthetic split documented in the dataset record. The transparency regimes ask about it, and reconstructing the split after the fact is not usually possible.

How we handle consent and licensing →

Related compliance topics

Not legal advice

We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.

Sourcing data under Synthetic Data?

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com