Getting voice actor consent for AI training, in practice
The briefing, the scope decisions, the withdrawal mechanics — and why the consent form is the least important part of the process.
Professional voice actors are asked to consent to AI training more often every year, and they have learned to ask precise questions. A process that treats consent as a signature collected at the end of a session produces one of two outcomes: a refusal, or a signature the actor did not really understand. An uninformed signature is worth very little, legally or ethically.
The work lives in two places: the briefing before the session, and the scope decisions made with the actor rather than presented to them.
The briefing, before anything is recorded
What has to be said out loud, in plain terms, and confirmed before the session starts:
- What the recordings will train — speech recognition, a synthetic voice, or both — described concretely rather than as "AI development."
- Whether a synthetic version of the actor's voice may be generated, and whether it may be used commercially.
- Whether the actor's name will ever be associated with the model or the product.
- Who will hold the recordings, and whether they may be shared, sublicensed or resold.
- How to reach a named contact with questions, now and later.
Scope is a set of separate decisions
Consent is not one yes. Each of these is negotiated and compensated separately, and each one is a place where a fair deal and an unfair one look identical on a one-page form:
- Term: perpetual, or a fixed number of years with renewal.
- Territory: worldwide, or limited to named markets.
- Use: one product, one model family, or any model built in the future.
- Exclusivity: whether the same voice may be licensed to others, and in which categories.
- Derivatives: whether the synthetic voice may be modified, and what content it may be made to say.
Withdrawal has to mean something achievable
Removing a single voice from the weights of a trained model is not a solved technical operation, and an agreement that promises it promises something nobody can reliably deliver. The remedy clause should list steps that can actually be performed:
- Stop using the recordings in future training runs.
- Delete the raw audio and derived files within a stated period, and confirm it in writing.
- Remove the voice from voice libraries, demos and marketing material.
- Where the voice exists as a standalone synthetic asset, retire it.
Record the process, not just the signature
What to keep, tied together by a session record: the signed form with the version number of the terms; the version of the briefing script; the date and format of the briefing; and, if the actor agrees, a short recording of them confirming the key points in their own words. The session record ties each delivered batch to the consent version that covered it.
The version history matters more than it looks. When terms are revised for a later project, the history shows precisely which actors agreed to which version — and prevents the quiet, dangerous assumption that an old signature covers new terms.
The extension problem
Most consent covers the project described at the time. A year later, when the model expands to a new use or a new market, the original consent does not stretch to cover it. There are two honest options: return to the actor for a new consent, which is slower but clean, or negotiate an extension clause at the start — a pre-agreed process and compensation for future scopes, defined in categories.
The extension clause is cheaper, and it only works when the actor trusts the process. That trust is why the clause has to be offered in the first briefing, before there is anything to gain from saying yes. The derivatives question, in particular, will produce most of the disputes in this field: an actor may happily consent to a navigation assistant and object to the same voice reading political advertising. Write the prohibited-content list into the agreement, in categories the actor helps define.