How to use a de-esser for smoother podcast speech

Harsh speech can make an otherwise polished podcast tiring to hear. The problem usually appears as sharp bursts around “s”, “sh”, “ch”, and “z” sounds, especially when a presenter moves close to the microphone or raises their energy for an important point. In audio production, this is called sibilance. A de-esser is designed to control those bursts without dulling the entire voice.

The best results come from treating a de-esser as a precise dynamic processor rather than a general-purpose tone control. It listens for excessive high-frequency energy, briefly reduces that selected range, and then lets the rest of the voice pass normally. Used carefully, it can soften a piercing consonant while preserving diction, breath, warmth, and personality.

This matters for Australian podcasts recorded in bedrooms, offices, community studios, and professional facilities alike. A bright microphone, reflective room, strong Australian “s” sound, or loud delivery can make speech uncomfortable on earbuds during a Sydney train commute or through a smart speaker in a Melbourne kitchen. Good de-essing keeps the programme clear across those everyday listening situations.

Find the harshness before processing

Listen for the problem in context rather than deciding that every “s” needs treatment. Sibilance often becomes obvious on words such as “series”, “business”, “Australia”, and “listeners”, but the harshest sounds may occur when a host becomes animated or leans towards the microphone. Solo the vocal briefly if necessary, then return to the full episode so you can judge whether the consonants are actually distracting.

A spectrum analyser can help locate the problem, although it should not replace your ears. Many speech recordings show troublesome energy somewhere between roughly 4 kHz and 10 kHz. Lower, thicker consonants may be concentrated near 3–6 kHz, while a narrow, airy hiss can sit higher. The exact range depends on the speaker, microphone, room, accent, and recording chain.

Do not confuse sibilance with unrelated problems. A noisy preamp, air conditioner, computer fan, clipped recording, or excessive room reflection will not be fixed by a de-esser. Nor will a microphone that is simply too bright become natural if its entire top end is being forced down. If the harshness is present on every word, consider microphone placement, acoustic treatment, equalisation, or a new recording before reaching for heavy processing.

Prepare the voice and signal chain

Record cleanly with sensible headroom. A spoken-word track that peaks around -12 to -6 dBFS gives you room for enthusiastic phrases and later processing. Avoid recording too quietly and then applying excessive gain, because raising the level also raises room noise. Equally, a clipped “s” cannot be restored convincingly by a de-esser.

Microphone position is an important first control. Placing the microphone slightly to the side of the mouth, rather than directly in front, can reduce the force of consonant bursts. A pop filter, a stable distance of about 10–15 centimetres, and consistent posture also help. The right position depends on the microphone, so move gradually while monitoring with headphones instead of relying on a fixed rule.

Use corrective editing before compression when possible. Remove obvious clicks, shorten distracting breaths only when they interrupt the rhythm, and reduce isolated noises manually. Compression can bring quiet sibilants forward, so a common chain is corrective editing, gentle de-essing, compression, equalisation, and then a final loudness check. Some engineers prefer to de-ess after compression if the compressor exaggerates the consonants; both approaches are valid when judged by ear.

For a multi-person Australian show, check each voice separately. A guest joining from Brisbane may sound darker than a host in Adelaide using a bright USB microphone, while a remote connection can add codec grit around the upper midrange. One de-esser setting across all speakers may over-process one voice and leave another sharp. Individual clip gain and track processing usually produce a more even result.

Set the de-esser controls carefully

Most de-essers provide a frequency or detection control, a threshold, a range or reduction amount, and an attack and release response. Some use a split-band design, which turns down only the selected high-frequency band. Others use broadband processing, which lowers the whole signal briefly when sibilance is detected. Split-band processing often preserves body, while broadband mode can sound smoother on voices with very wide, aggressive consonants.

Begin with the detector frequency. Play several harsh words and sweep the control until the de-esser responds to the hiss rather than vowels or general brightness. Then lower the threshold until only the strongest sibilants trigger gain reduction. A useful starting point is around 2–4 dB of reduction on normal peaks, with perhaps 5–6 dB on unusually sharp words. The correct amount is the least processing that makes the voice comfortable.

Attack should be quick enough to catch the beginning of the consonant, but an extremely fast setting can make diction feel lisped or soft. Release should return the high end naturally after the sound passes. If the release is too slow, the next vowel may become dull; if it is too quick, the processor can chatter on sustained “s” sounds. Adjust while listening to sentences, not isolated test tones.

Many plug-ins offer a listen or audition mode that isolates the detected material. Use it briefly. The isolated signal should mostly contain “s”, “sh”, and related consonants, not the whole vocal or prominent vowel tones. If the audition contains a large amount of the voice, refine the frequency range or use a narrower detection filter before increasing reduction.

Combine automatic control with manual editing

A de-esser is efficient for a long interview, but it does not need to do every job. One unusually sharp word can cause a processor to reduce the high end of an entire phrase. Clip gain, region gain, or volume automation can lower that consonant before the plug-in sees it. This keeps the overall setting gentle and prevents audible pumping.

For a recurring presenter, save a starting preset but do not apply it blindly. Mic distance, guest mood, room position, and different recording days all change the result. Compare the processed and unprocessed voice at matched loudness; a louder version can seem clearer even when it is actually harsher. Bypass the plug-in at intervals to make sure the voice still sounds like the same person.

Keep the natural rhythm of speech. Removing every audible “s” makes narration unclear, particularly for listeners using a phone speaker or inexpensive earbuds. Australian place names and terms can also lose their character if consonants are softened too aggressively. The target is reduced fatigue, not a lisp, whisper, or artificially dark broadcast voice.

For a carefully edited episode, create notes about troublesome sections and hand them to the mixing engineer. Clear arrangement and production notes help everyone work consistently, particularly when an episode includes music, advertisements, multiple speakers, or a recurring segment. The guidance in song arrangement notes is written for music, but its principle of communicating precise timing and intent is equally useful when marking spoken-word edits.

Check the result across Australian listening conditions

Podcast listeners rarely use one monitoring system. Check the voice on studio headphones, ordinary earbuds, a laptop, and a small mono speaker. A de-essed track that sounds pleasant on large monitors may still be piercing on AirPods, while a track that sounds smooth in headphones may become muddy on a phone. Listen at a modest volume as well as at the louder level used for detailed editing.

Australian podcast audiences may listen during driving, walking, housework, commuting, or outdoor exercise. Strong background noise can mask low-level detail, encouraging producers to brighten the entire vocal; that often makes sibilance worse. A balanced midrange, stable loudness, and controlled high frequencies usually translate more reliably than an aggressively bright master.

If the podcast contains advertising, compare the presenter’s voice with inserted sponsor reads. Ad copy is often delivered with extra energy, and a bright commercial recording can make the main conversation seem dull by comparison. Process the voices so that their tonal balance and loudness feel coherent, while preserving the distinction between editorial speech and paid content.

There are also compliance considerations beyond sound. If guests or contributors can be identified, obtain appropriate consent and explain how their recording will be used. Australian privacy obligations can apply to the collection and publication of personal information, and state or territory rules about recording private conversations vary. Keep release forms, usage permissions, and episode notes organised rather than assuming that a technically clean recording is automatically cleared for publication.

Finalise the episode without over-processing

After de-essing, apply compression with restraint. A ratio around 2:1 or 3:1, a moderate attack, and a release that follows the speaker’s pace can even out level without flattening expression. If compression brings new hiss forward, use a second very light de-essing stage rather than making the first processor work excessively. Two subtle stages can sound more natural than one severe stage.

Use equalisation to support intelligibility, not to compensate for poor de-essing. A gentle high-mid adjustment may improve clarity, while a small low-cut can remove rumble from traffic, air conditioning, or desk vibration. Be cautious with broad boosts above 6 kHz, because they can make already-sibilant speech more abrasive. Automate level differences between speakers before chasing them with extreme compression.

When you are ready to export, check the entire episode from beginning to end. Listen for words that sound lisped, sudden changes in brightness, missing consonants, or breaths that become unnaturally loud after processing. Confirm that music does not mask the speech and that transitions do not create a sharp jump in perceived volume. Leave a clean archive of the edited tracks and processing notes so a later revision does not require rebuilding the episode.

Practical checks for a controlled vocal

  • Locate the main sibilance range with your ears and a spectrum analyser, then adjust the detector to that area.
  • Start with gentle gain reduction and increase it only when harsh consonants remain distracting.
  • Compare split-band and broadband modes to see which preserves the speaker’s natural tone.
  • Use clip gain for isolated problem words instead of forcing the whole track through heavy processing.
  • Check the processed voice on headphones, earbuds, a laptop, and a small mono speaker.
  • Review sponsor reads, remote guests, and host dialogue together for consistent loudness and brightness.
  • Keep consent records, usage permissions, and the unprocessed session files with the finished episode.

A professional recording environment can make this work easier by providing controlled acoustics, suitable microphones, accurate monitoring, and an experienced second opinion. For podcasts recorded in Elgin, Illinois, or prepared remotely for listeners across Australia, careful de-essing remains a small but important part of a reliable spoken-word workflow. The finished voice should feel close, intelligible, and comfortable without drawing attention to the processing itself.

20
Simultaneous Channels
15+
Microphones
5
Featured Artists
30+
Years PC Experience
Get Started

Book Your Session

Ready to record? Reach out to discuss your project, check availability, and get answers about the studio setup.

1
Tell us about your projectShare details about your band, genre, and what you're looking to accomplish.
2
Discuss availability & ratesWe'll go over scheduling options and affordable hourly pricing.
3
Come in and recordWork in a comfortable studio with pro-grade gear and an experienced engineer.
Send an Inquiry
Why LnL Recording

Built for Serious Sound

A studio designed around a custom PC platform — not an off-the-shelf solution — with decades of hands-on engineering experience behind every session.

⚡
Custom DAW
A purpose-built PC workstation configured specifically for multi-track audio production, not a generic computer.
🎤
Pro Mic Locker
AKG C414B-TLII, Shure KSM32/SL, SM57, SM81, Sennheiser e906, AKG D112, and many more — the right mic for every source.
🎸
Instruments & Amps
Gibson Les Paul Deluxe, Ibanez 540S, Marshall JVM410, Vox AD100VT, Korg Triton Pro, Kurzweil SP88, and a full drum kit.
The Equipment

Studio Gear

A comprehensive collection of microphones, preamps, outboard processing, instruments, and monitoring — all detailed on the Studio Gear page.

Microphones
  • AKG C414B-TLII — Large-diaphragm condenser
  • Shure KSM32/SL — Studio condenser
  • Shure SM57 & SM48 — Dynamic workhorses
  • Shure SM81, AKG C1000S — Small-diaphragm condensers
  • CAD Equitek E-100 — Supercardioid condenser
  • Oktava MC012 — Multi-capsule condenser
  • Sennheiser e906, E602 — Guitar cab & kick drum
  • AKG D112, EV N/D 468, Audix Fusion 6
  • Studio Projects C1
Preamps & Processing
  • Mackie MS1642-VLZ4 mixer
  • Black Lion Audio Auteur Quad preamp
  • Black Lion Audio B173 preamp
  • Joe Meek VC1QCS channel strip
  • ART TubeMP tube preamp
  • Radial J48 active DI
  • Rocktron Intellifex effects processor
  • BBE 462 Sonic Maximizer
  • Rolls RA62HA — 6-output headphone amp
Full Gear List
Past Sessions

Client Roster

Artists and bands who have recorded at LnL Recording.

SIIN
Dead Man's Hand
Sthica
Sean Begora
The Olsen & Pahl Project
View Client Roster
Let's Work Together

Ready to Record?

Located in Elgin, Illinois. Reach out to discuss your next project.

Contact Us