Voice cloning: how much audio, what quality, and where consent ends
Cloning needs one speaker recorded consistently rather than many speakers recorded widely. The hard parts are coverage, session drift, and consent that answers what happens to the model.
Voice cloning builds a synthetic voice that sounds like one specific person. Unlike almost every other speech project, the speaker count is one, and the question of how much audio is needed has a different answer depending on what the voice has to do.
That dependency is worth stating first, because it is the most common source of a failed purchase. A buyer who asks for "enough audio to clone a voice" receives a number that was right for a different target: a narration voice that reads in one style needs far less than a conversational voice that has to handle interruption, laughter, and emotional range.
The amount depends on the target, so specify the target
Three targets cover most projects, and each one implies a different recording plan.
A single-style voice — read-aloud narration, announcements, a fixed set of prompts — can be built from a small amount of audio. The commonly used range for adaptation-style approaches starts at tens of minutes and extends to a few hours, with returns flattening quickly once the style is consistent.
A voice that must speak arbitrary text naturally is a coverage problem rather than a duration problem. The text it will be asked to say includes sound combinations and prosodic patterns that may be rare in a small set, so the script is chosen for phonetic and prosodic coverage: a balanced sentence set rather than pages of prose.
A voice that has to be expressive — pleased, annoyed, apologetic, surprised — needs those styles recorded as styles. A model trained only on neutral reading produces a flat voice on emotional text, and more neutral audio does not fix it. Style coverage is its own line in the plan, and it is the one most often left out.
Consistency is the quality bar, not fidelity
The characteristic failure of a cloned voice is instability: the same sentence rendered twice comes out with a slightly different character, so a long passage sounds like several people taking turns. That usually traces back to the recording rather than the model.
The causes are mundane and controllable. A different microphone, a different room, a different distance from the microphone, a different time of day, a different level of vocal effort after a long session. Each one introduces a variation the model learns as part of the voice.
So the recording plan should pin the chain: one microphone, one room, one distance, one posture, and sessions short enough that fatigue does not change the voice. Long sessions should be split, and each one should be checked against the first for drift — the same reference sentence recorded at the start and end of every session is the cheapest possible control.
Record above the target sample rate of the synthesizer and reduce later; the reverse is not possible. And check the boring problems before the interesting ones: clipping, room hum, and background noise that changes with the hour of the day.
Consent is the foundation of the asset, not a formality
Cloning is the case where consent carries the most weight, because the thing being created is a person's identity. A consent that only covers "being recorded" is not sufficient for this use, and the gaps tend to surface later as disputes rather than as clean legal findings.
The agreement has to answer, in plain terms: that the recordings will be used to build a synthetic voice; what that voice may be used to say, and whether the speaker has any right to see or refuse content they have not been shown; how long the permission lasts; whether the resulting model counts as a deliverable and who holds it; and what withdrawal means — including what happens to a voice model that has already been trained, which is not obviously reversible and should be answered before recording rather than after.
Two cases need their own process rather than a variation of the standard one: a child's voice, and the voice of someone who has died. Both involve a person other than the speaker holding the decision, and both attract scrutiny that ordinary recording work does not.
One rule holds regardless of what any document says: a cloned voice must not be used to imitate a third party. A speaker can authorize their own voice; they cannot authorize someone else's.
Why public recordings do not work
The shortcut is tempting because the audio already exists: podcasts, interviews, lectures, videos. It fails for three independent reasons, and any one of them is enough.
Rights first. A published recording is protected, and permission to listen to it, or even to quote it, is not permission to train a model on it. The speaker's agreement is a separate question from the recording's copyright, and both are required.
Technical second. Scraped audio is mixed with music, compressed, recorded on unknown equipment at unknown distances, and often contains more than one speaker. A model trained on it learns those conditions as part of the voice, and the result degrades on any text that does not resemble the source material.
Identity third. The output is a copy of a real person, made without their agreement, used to say things they never said. That is the scenario the rules are converging against, and it does not become acceptable because the audio happened to be public.
What to put in the recording agreement
A session that fails quality control should be replaced rather than compensated for, which is why the last item on this list is worth arguing about before the first session rather than after the tenth.
- One speaker, with the session count and the script coverage list attached.
- The fixed recording chain, so that every session is comparable to the others.
- The drift check: the reference sentence at the start and end of each session, and the threshold for rejecting a session.
- The quality criteria on delivery: noise floor, clipping, and consistency across sessions.
- The consent scope, its term, and a written answer to what happens on withdrawal.
- The re-recording clause, so a failed session is replaced rather than tolerated.