Cross-border transfers of training data: what counts as a transfer, and how to structure one

A training corpus crosses borders more than once — collection, annotation, storage, compute — and each of those moves is its own legal event. Here is how to map them and what to write down.

A corpus crosses borders more often than anyone counts

Ask a project team how many countries their data touches and the answer is usually one: where the speakers were. The realistic count is higher, because every stage of the pipeline can sit somewhere different, and each one is a separate event with its own paperwork.

The events to map individually:

  • Collection: where the recording happens and where the first copy is stored.
  • Annotation: where the people who listen and transcribe are located — and freelance annotators are located wherever they are, not where the vendor's office is.
  • Storage and backup: where the files live and where the disaster-recovery copy lives, which is frequently a different country.
  • Compute: where the training run executes, which can be a fourth jurisdiction again.

Remote access counts as a transfer

The point that decides most disputes: an annotator in a third country who streams audio from a server in the EU is receiving a transfer, even if no file ever lands on their machine. Access is processing, and processing from outside the region is an export regardless of where the bytes are stored.

The same logic catches the tooling. A web-based annotation interface that caches audio locally on the annotator's device creates copies in jurisdictions nobody planned for, and those copies persist after the session ends unless the tool is configured not to keep them. A pipeline diagram that shows vendors but not tools will understate the number of countries involved.

A useful exercise: instead of drawing the data flow, list every person and system that can play a recording, and write down where each one sits. That list is the transfer map, and it is usually longer than the vendor list.

The mechanisms, and the traps inside them

Under the GDPR a transfer outside the EU needs a legal mechanism in place before it happens. Three carry most training data work. An adequacy decision treats the destination country as providing equivalent protection; it requires the least paperwork and it can be withdrawn, so it is a mechanism to track rather than to rely on indefinitely. The standard contractual clauses are the Commission's modular terms, and the module that applies depends on whether each party is acting as a controller or a processor for that flow — choosing the wrong module is a common and entirely avoidable defect. Derogations cover occasional, limited transfers: useful for a one-off, not a structure for a recurring pipeline, and not something to build a data business on.

Two traps sit inside those mechanisms, and both are easy to walk into. The first is that an adequacy decision for a country does not cover every company in that country. Under the framework for the United States, only organisations that have certified under it are covered, so the importer's status has to be checked against the official list rather than assumed from the country. A supplier that has not certified is a supplier for which some other mechanism is needed, even though the country itself is treated as adequate.

The second trap is treating a signature as the end of the analysis. The parties are expected to assess whether the law of the destination undermines the protection the clauses promise, and to add measures where it does. For a training corpus, that assessment has to consider government access to data held by the importer — a question about the importer's legal environment rather than about the contract.

Both traps have the same practical answer: name the mechanism per flow, name the importer's status, and revisit both on a schedule, because adequacy decisions and certification lists change without notice to the parties relying on them.

Where speech data makes it harder

Audio adds complications that a text pipeline does not have:

  • Listening is access. Everyone who can play a recording is processing it, which multiplies both the locations to map and the subprocessors to name.
  • Annotation tools cache, and cached audio is the copy that escapes the pipeline diagram.
  • Voice is treated as sensitive or biometric in some destinations even where the source country does not classify it that way, which changes the assessment even when the mechanism is unchanged.
  • Quality review is often done by a second team in a third location, which is another transfer nobody listed at scoping.

Design choices, and what to write down

Some transfer exposure is designed away rather than papered over: keep raw audio in one region and ship only what each stage needs, such as transcripts for a text pass and short segments for spot checks. Require annotation platforms that do not persist audio locally, or audit the ones that do and record what you found. Store audio separately from identifying metadata, so the more sensitive artifact has fewer copies in fewer places. And fix the training region in the contract rather than leaving it to whoever schedules the job.

The clauses that make a transfer file survive a review:

  • The named mechanism per data flow, rather than one mechanism described once for the whole relationship.
  • The full list of processing locations, including parties that only access rather than store.
  • A change-notice obligation before a new location is added, with a right to object within a stated period.
  • The importer's status under any adequacy framework, warranted and re-warranted on a schedule.
  • An allocation of who bears the cost if a mechanism is invalidated mid-project — the clause nobody wants to negotiate and everybody eventually needs.

The short version

Transfers in a data business are a mapping problem before they are a legal one. Count the locations, including the ones that only listen; pick a mechanism that fits the roles; check the importer rather than the country; and put the answers in the contract instead of in a project email.

This is an operational overview rather than legal advice. Transfer law changes by decision and by case, and the right structure depends on the roles and countries involved in a specific project.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com