Drafting the DPA schedule for a speech data project

A speech data DPA schedule needs different details than a text one: listening access, audio retention, and subprocessors who only listen.

The main body of a data processing agreement is fairly standard across industries. The schedule is where the project actually lives, and for speech data a schedule copied from a text-data template will miss the parts that matter: who is allowed to listen, how long raw audio survives, and which subcontractors count as subprocessors.

What follows is a structure to fill in for a speech data project, section by section — not a template to sign, because the answers differ by project and the schedule is where those differences have to appear.

Part A: the processing description

The description should be specific enough that a reader who has never seen the project can picture it:

  • Nature of processing: recording, segmentation, transcription, diarization, quality review, packaging.
  • Purpose: preparing training, validation and test data for named model families, not "AI development" in general.
  • Data categories: voice recordings, transcripts, speaker metadata, consent records — with a flag on any category treated as sensitive under the applicable regime.
  • Data subjects: paid session speakers, field participants, and any bystanders audible in the files.
  • Duration of processing, and the specific event that ends it.

Part B: technical and organizational measures, speech edition

The measures section is where a text template shows its limits. Speech adds a category of access — the ability to hear the data — that has no clean equivalent in text:

  • Listening access is a named, logged permission: who can play audio, on which systems, with what record.
  • Raw audio is encrypted at rest and in transit, and access to audio is separated from access to transcripts.
  • Annotation environments that cache audio locally are prohibited or audited; the local cache is where copies escape.
  • Speaker-identifying metadata is stored separately from the audio wherever the workflow allows.
  • Deletion of audio is verified across primary storage, backups and tool caches, with written confirmation.

Part C: retention, with separate clocks

Speech workflows run at least four retention clocks at once, and the schedule should set each one deliberately:

  • Raw audio: the shortest period that still allows quality disputes and re-annotation, with a defined trigger — usually acceptance of the delivery.
  • Transcripts: often kept longer, and often less sensitive than the audio they came from.
  • Consent records: must outlive the data they justify, with their own retention statement.
  • Quality samples: a small, deliberately chosen retained set for future reference, not accidental leftovers.

Part D: subprocessors, including the ones who only listen

Subprocessor categories to list for a speech project:

  • Annotation platforms and the freelance annotators who work on them — individuals count, not just companies.
  • Transcription and translation vendors.
  • Quality review and calibration contractors.
  • Cloud storage and compute providers.
  • Any party that can stream or play the audio, even without downloading it.

Keeping the schedule alive

Deletion deserves its own sentence in the schedule, because audio is harder to delete than text. Every player, annotation tool and backup holds a copy, and a retention clause that says "deleted after acceptance" without saying how deletion is verified will not survive its first audit. Streaming playback in a browser is processing, so a subcontractor who never holds a file but hears every recording belongs on the subprocessor list, listed by category with a change-notice process rather than by named individuals who go stale within a quarter.

Review triggers to write into the agreement: a new tool in the annotation pipeline, a new supplier, a change in where data is stored, a change in what the model is trained for, and any breach or near miss. Assign one owner on each side, and make the review a contractual obligation rather than a good habit. In a dispute, the schedule that applies is the one in force when the processing happened — so version control on this document matters as much as its content.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com