Planning audio recording sessions so they produce usable hours
Session length, fatigue, and in-session checks decide how much of a recording day survives review. A plan for the day, and the numbers to set it with.
The session is where the cost is decided
A recording day costs the same whether it produces four hours of accepted audio or four hours of material that fails review. Everything that decides which of those happens is settled inside the session: the length of the blocks, the state of the speaker, the state of the room, and whether anyone checked the written file rather than the live feed.
A session plan is therefore not a schedule for the speakers. It is a quality control plan with speakers in it.
How long a session should be
The limit is not the clock. Vocal fatigue and attention degrade output well before a speaker is unwilling to continue, and the degradation is gradual, which is what makes it dangerous: the material still sounds fine and it scores worse.
A shape that holds up for read speech: blocks of forty-five to sixty minutes of recording separated by ten to fifteen minute breaks, three to four blocks in a day, and a hard stop that is respected even when the speaker offers to continue. Useful output per speaker per day for read material typically lands in the low single digits of hours, and the second half of a long day is measurably worse than the first.
- Read speech is limited by vocal load and by attention to the text. Speakers speed up as they tire, which changes the very property a read corpus is supposed to hold constant.
- Spontaneous and conversational speech is limited by cognitive load and topic exhaustion. The third hour of describe what you did yesterday produces repetition, and repetition is what makes the extra hours worthless.
- Sustained vocal tasks, such as phonetically balanced lists or held vowels, are physically demanding. They belong in short blocks, and not at the end of a long day.
- A new speaker's first session is shorter than a returning speaker's. Part of it goes to setup, orientation, and a warm-up that gets discarded.
The first ten minutes are not data
Speakers over-articulate when they are being careful, so the opening minutes of a session are the least representative of how that person actually speaks. Plan a warm-up of reading or free speech that is discarded, and start the retained recording once the speaker has settled.
The same effect appears at the other end. Fatigue shows up as a rising speaking rate, falling intensity, more creak in the voice, a flattening pitch range, and more disfluency. A session lead who knows the signs can end a block early; one who does not will record an hour of material that fails review.
Check the file, not the monitor
Live monitoring on headphones catches a bad take. It does not catch a file written at the wrong rate, with a dropped channel, or with a header that claims something the content does not contain. These are two different checks and both are needed.
- Before each block: record ten seconds of silence and measure the noise floor. Rooms change during a day as traffic, ventilation, and neighbouring activity change.
- At the start of each block: a level check, because a speaker who leans back or turns away changes the input level even when nothing else has moved.
- During each block: a spot listen of thirty to sixty seconds played back from the written file, done by the session lead rather than the speaker. This is the check that catches configuration errors.
- At the end of each block: the file count against the expected count, and the duration against the planned duration, written into the session record while the session is still happening.
- Before the session starts: the room list. Fans, refrigeration, fluorescent ballasts, phone notifications, squeaking chairs, jewellery, and clothing that rustles against a microphone. Each of these has ruined a session that was otherwise well run.
The session record is written during the session
Everything the manifest needs is knowable only inside the session: speaker code, session code, date, room, equipment serial numbers, gain settings, script or prompt version, hours actually recorded, and every deviation from the plan.
Deviations recorded at the time are documented variation. The same deviations remembered a week later are unexplained defects, and the difference between the two is one line in a log. Record the deviations the producer would rather not mention, because those are the ones that surface in review.
Rules that apply across the schedule
The rules below operate at the level of the schedule rather than the session, and each one exists because the problem it prevents is invisible on any single day.
- Cap the number of first sessions per session lead per day. A new speaker needs more attention, and attention does not scale with the number of rooms.
- Keep a buffer day per week for re-sessions. A re-session cannot be batched with scheduled speakers, so it consumes disproportionate calendar time when it is squeezed in.
- Do not schedule the most demanding task for the last block of the day, for any speaker.
- Track the re-session rate by reason code and treat it as a process metric. When the rate rises, the cause is usually a misconfigured recorder, an unchecked room, or a speaker who was mis-screened, and all three are process failures rather than speaker failures.
- Do not put the same speaker in two sessions in one day unless the protocol requires it. Where it is unavoidable, separate them by the longest break the schedule allows, and compare the second session's first block against the first session's last.