Best practices for recording voice-over for an explainer video
Voice-over carries the story in an explainer video. Even when the animation is crisp and the music is well-scored, a weak or noisy narration will sink the whole piece. The voice is the through-line that guides viewers from one idea to the next, and explainer videos tend to live or die on how clearly that voice is captured, edited, and delivered.
In Australia, the explainer market is unusually busy. Sydney and Melbourne host a steady stream of agencies producing content for fintech startups, mining services, government departments, and tourism boards, while Brisbane and Perth studios are catching up fast as remote production pipelines mature. Many Australian voice-over artists work from home booths that range from professionally treated rooms to closet setups, and clients often expect broadcast-quality audio without the budget of a network production. That mix of high expectations and modest budgets makes disciplined technique more important than ever.
Preparing the script for a clean read
A voice-over session begins long before the microphone is switched on. The script needs to be edited for the ear rather than the eye. Long sentences should be broken up, technical jargon either rewritten or briefly defined, and every comma should mark a real breath or thought pause. Phrases that read fine on paper can feel clumsy when spoken, so a quick pass aloud often catches tongue-twisters before they reach the booth. Marking up the copy with underlining for emphasis, slashes for pauses, and notes on energy levels gives the performer a roadmap and saves takes later.
Pronunciation is a particular trap when the script mixes international terms with Australian English. Words like "data", "process", and "container" sound slightly different in Australian English than in American English, and place names from overseas need consistent handling. For projects aimed at a domestic audience in Sydney, Melbourne, or regional centres like Adelaide, the natural local rhythm reads as authentic, but the brief should still spell out how brand names, acronyms, and product terms should sound. Local industry conventions in Australia often favour a relaxed, conversational delivery over a hard-sell tone, and that preference should be reflected in the script's punctuation and paragraph breaks.
Rehearsal matters as much as mic technique. Most experienced narrators will run the script cold three or four times before any recording begins, paying attention to where their energy naturally rises and falls. A short warm-up of tongue twisters and sustained vowels loosens the mouth and steadies the breath. The aim is to arrive at the session with a settled pace, so that the time in front of the microphone is spent on interpretation rather than basic fluency.
Choosing the right room and microphone
The room matters more than the microphone for spoken-word recording. Hard, parallel surfaces produce flutter echo and comb filtering, and the small reverberant tail that sounds acceptable in a music session can make a voice-over sound distant and indistinct. A well-treated booth with broadband absorption on the rear wall, side panels to kill early reflections, and a thick carpet or rug underfoot will deliver a dry, intimate read. For artists working from home, even a walk-in wardrobe packed with clothes and a couple of moving blankets can produce usable results for short explainer pieces.
Microphone choice for narration usually points to a large-diaphragm condenser with a cardioid pattern. Models with a smooth, slightly warm top end tend to flatter spoken word without requiring heavy equalisation later. A pop filter sits a few centimetres in front of the capsule, and a shock mount isolates the mic from vibrations travelling through the desk or stand. Distance from the mic is a creative choice as well as a technical one: closer placement adds intimacy and bass proximity, while working 20 to 25 centimetres back keeps the tone more neutral and leaves room for the performer to lean in for emphasis.
Gain staging is just as important here as it is in louder music sessions, and the same care that goes into recording loud rock bands cleanly applies in reverse. The voice peaks should sit comfortably below 0 dBFS on the DAW meter, ideally around -12 to -6 dBFS, with healthy headroom for any unexpected loudness from a laugh, a sigh, or a punchy consonant. Watching the levels as carefully as a band engineer watches a snare drum prevents distortion and keeps the noise floor low when the compressor later pulls quiet passages upward.
Performance, pacing, and tone
Explainer voice-overs live in a narrow band of tone. Too formal and the viewer tunes out; too casual and the message loses authority. The trick is to imagine speaking to a single smart colleague who has just asked, "So how does this work?" That mental picture produces a friendly, confident delivery without slipping into either a corporate drone or a parody of casualness. Many Australian clients expect this conversational warmth, especially for consumer-facing brands in retail, health, and tech.
Pacing should match the on-screen edit. If the animation cuts every four to five seconds, the narration needs to deliver one clean idea per cut. Reading too slowly leaves dead air and makes the visuals feel sluggish; reading too quickly overwhelms the viewer and forces the editor to either speed up the animation or cut sentences. A simple rule is to time the script at the intended pace, then aim to record slightly under that mark to leave room for natural pauses at the end of sentences and at paragraph breaks.
Breath control is the silent partner of good pacing. Narrators who run out of air halfway through a sentence tend to speed up the tail of the line, which registers as anxiety in the listener. A short pause before a difficult phrase, an unhurried exhalation through the nose, and a sip of room-temperature water between sections all keep the voice steady. Recording a single, fully committed take of an entire script is rarely practical, but aiming for that level of focus on every paragraph pays off. The discipline of recording a band live shares a similar philosophy: commitment to the moment lifts the read above a stitched-together assembly.
Editing and processing the take
Once the session is wrapped, the real assembly begins. The first task is to comp the best sections from multiple takes into a single, flawless read. Editors listen for clean word starts, smooth transitions between sentences, and consistent energy across paragraphs. Software tools that allow playlist-style comping make this faster, but the ear still has to be the final judge. A short gap of silence between sentences should be long enough to feel natural but short enough to maintain momentum, usually somewhere between 200 and 400 milliseconds.
Cleanup comes next. Mouth clicks, lip smacks, and hard breaths are easy to spot on a waveform and quick to remove with a spectral editor or a dedicated de-click plugin. Background hum from lighting, computer monitors, or air conditioning is dealt with by a noise profile captured during a quiet moment, usually at the start or end of the session. The goal is not to make the voice sound processed but to make the issues invisible. Over-processing leaves artefacts that listeners may not name but will certainly feel.
Subtle processing rounds out the voice. A high-pass filter around 80 Hz removes rumble without thinning the tone. Gentle compression, usually no more than 3:1 with a slow attack and a medium release, evens out the dynamic range and helps the voice sit on top of music and sound design. A de-esser tames any harshness on sibilant consonants without dulling the rest of the voice. EQ is usually conservative: a small dip in the lower mids around 250 Hz can clear muddiness, while a gentle lift near 5 kHz adds clarity without sharpening sibilance. The processed voice should sound like the performer on a good day, not a redesigned instrument.
Delivering the final audio
The final file format depends on the downstream workflow, but most explainer producers expect a high-resolution WAV or AIFF as a master, with an MP3 supplied for quick review. A 48 kHz sample rate at 24-bit depth is the modern broadcast and web standard and gives plenty of headroom for any processing that may happen further down the chain. The sample rate should match the project, not be changed at the last minute, as sample-rate conversion can shift timing in subtle but audible ways.
Loudness targets vary by platform. Australian broadcasters follow standards aligned with international broadcast norms, but explainer videos for the web typically aim for an integrated loudness around -16 to -14 LUFS, with a true peak ceiling of -1 dBTP to avoid clipping on consumer devices. Checking the final mix against these targets with a loudness meter saves time in revisions and protects the client from playback issues on phones, laptops, and TVs.
The delivery should include the raw, unprocessed takes alongside the edited master. Many Australian producers appreciate having the option to remix or re-time the narration if the animation changes, and that backup is worth its weight in late-night project rescues. File naming should follow the project's convention, with version numbers, take labels, and a brief readme explaining the contents. Clear delivery turns a good voice-over into a reliable building block for the rest of the production, and that reliability is what keeps explainer studios in Sydney, Melbourne, and beyond coming back for another round.
Book Your Session
Ready to record? Reach out to discuss your project, check availability, and get answers about the studio setup.
Studio Services
Professional multi-track recording with world-class gear at affordable hourly rates.
Built for Serious Sound
A studio designed around a custom PC platform — not an off-the-shelf solution — with decades of hands-on engineering experience behind every session.
Studio Gear
A comprehensive collection of microphones, preamps, outboard processing, instruments, and monitoring — all detailed on the Studio Gear page.
- AKG C414B-TLII — Large-diaphragm condenser
- Shure KSM32/SL — Studio condenser
- Shure SM57 & SM48 — Dynamic workhorses
- Shure SM81, AKG C1000S — Small-diaphragm condensers
- CAD Equitek E-100 — Supercardioid condenser
- Oktava MC012 — Multi-capsule condenser
- Sennheiser e906, E602 — Guitar cab & kick drum
- AKG D112, EV N/D 468, Audix Fusion 6
- Studio Projects C1
- Mackie MS1642-VLZ4 mixer
- Black Lion Audio Auteur Quad preamp
- Black Lion Audio B173 preamp
- Joe Meek VC1QCS channel strip
- ART TubeMP tube preamp
- Radial J48 active DI
- Rocktron Intellifex effects processor
- BBE 462 Sonic Maximizer
- Rolls RA62HA — 6-output headphone amp
Client Roster
Artists and bands who have recorded at LnL Recording.
Ready to Record?
Located in Elgin, Illinois. Reach out to discuss your next project.
Contact Us