Global

Ethical Data Collection

The short answer. Ethical data collection is the set of practices that go beyond the legal minimum: paying speakers at a rate that reflects the session, telling them what the model is for before they agree, not recording people in public without notice, and documenting the terms that were actually explained rather than the terms on the form. EU AI Act Article 10 gives the buyer a reason to care, because the provider has to document the origin of the data and the original purpose of collection. GDPR Article 13 requires the transparency in the first place.

The law

EU AI Act Article 10(2)(b)

data collection processes and the origin of data, and in the case of personal data, the original purpose of the data collection

The provider of a high-risk AI system has to document where the data came from and why it was originally collected. That record is written by the buyer, from facts supplied by the vendor, and it has to be accurate. If the original purpose on file is "speech technology research" and the actual use is a distributed commercial model, the record is wrong, and the vendor is the only party that can correct it.

EU AI Act Article 10(2)(f)

examination in view of possible biases that are likely to affect the health and safety of persons, have a negative impact on fundamental rights or lead to discrimination prohibited under Union law…

This is why the composition of a corpus is a governance question and not only a quality question. Who was recorded, who was left out, and how the speakers were recruited all bear on whether the dataset is suitable for its intended purpose. A corpus recruited entirely through one university, or one city, has a documented limitation, and the limitation is easier to state at purchase than to explain later.

GDPR Article 13(1)(c)

the purposes of the processing for which the personal data are intended as well as the legal basis for the processing.

The transparency duty: the speaker has to be told the purpose and the legal basis at the time of collection, in a form they can understand. This is where ethical practice and legal compliance meet. A form explained in a language the speaker does not read satisfies the paperwork and not the requirement, and consent that was not understood fails the validity test rather than merely the spirit of it.

Who it applies to

Article 10 of the AI Act binds providers of high-risk AI systems. It is not a rule addressed to data vendors, and the reason it belongs in a procurement document is that the buyer's obligation can only be discharged with information the vendor holds. The vendor's collection practices become a contractual dependency of the buyer's compliance, whether or not the vendor is a regulated party.

Article 13 binds whoever collects from the data subject, which is the vendor when it recruits speakers directly, or the buyer when it commissions a recording. The obligation follows the collection, not the delivery.

The practices described here apply regardless of jurisdiction. A corpus collected in a country with no data protection statute is still the corpus that has to be defended in a diligence review, and the speakers in it are still people whose cooperation the next project depends on. None of these practices is required by the law of every country where recordings happen, and all of them are cheaper than re-collecting a corpus after the recruitment method becomes a problem.

What it costs to get wrong

There is no fine for unethical collection as such. The costs arrive by other routes, and they are not smaller.

If the explanation was not understood, the consent is not informed, and the GDPR validity test fails. That moves the dataset into the higher tier under Article 83(5): up to 20 million euros, or up to 4% of total worldwide annual turnover of the preceding financial year, whichever is higher.

If the AI Act governance record cannot be substantiated, the provider's Article 10 compliance is in question, and Article 99(4) sets fines of up to 15 million euros or 3% of total worldwide annual turnover for non-compliance with provider obligations.

The reputational cost is the one buyers underrate. If the recruitment practice can be described in a single sentence that a journalist would publish, the cost attaches to the model rather than to the vendor, and it is not covered by an indemnity. The operational cost follows: a corpus that cannot be documented cannot be extended, because the speakers will not come back, and for uncommon languages there is no replacement pool to go to.

How to comply when you are buying data

These are the practices we look for when assessing a collection partner, written so they can be asked about directly.

  • Pay for the session, not the outcome. A fair rate is a session fee benchmarked to local professional rates for comparable time — not a per-clip micro-payment, and not a payment contingent on the data being accepted by the end buyer. Payment on acceptance creates pressure to perform that undermines whether the consent was freely given.
  • State the use in the form and in the explanation: that the voice will be used to train a model, that the model may be distributed, and who may receive it. If the honest answer is "a model that may be licensed to third parties", that sentence goes in the form.
  • Do not record in public spaces without notice. If ambient speech is captured, the people in it are data subjects, and a notice on a door is not consent from them.
  • Document what was explained, not only what was signed: the language used, who did the explaining, the date, and whether the speaker asked questions. Keep it with the consent form, because the explanation is what a validity challenge tests.
  • Pay within a stated period and the same amount whether or not the recording is used. Predictability is what makes the arrangement voluntary rather than conditional.
  • Ask the vendor for the recruitment method and the pay rate for the corpus you are buying. That answer tells you more about how the data will perform than the accuracy metrics do, because speakers who were treated well produce recordings that are worth keeping.

How we handle consent and licensing →

Related compliance topics

Not legal advice

We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.

Sourcing data under Ethical Data Collection?

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com