De-Essing Dialogue: Controlling Sibilance with Frequency-Based Dynamics
Sibilance is one of the fastest ways to tire out a viewer's ears in a dialogue mix. A field-tested approach to placing a de-esser, finding the right band, and choosing between manual clip gain and dynamic EQ when the plugin alone isn't enough.
On a documentary mix, every 's' from the narrator went into the ear like a blade. The director said 'the voice is too bright', but I knew brightness wasn't the issue — only a handful of consonants were out of control. Cutting the whole top end with EQ would have made the dialogue dull. The answer was a tool that acts only when sibilance happens: a de-esser.
What Is Sibilance and Where Does It Live?
Sibilance is the intense high-frequency energy in consonants like 's', 'sh', 'z' and 'ch'. It varies with the speaker, the microphone and the language, but the energy typically clusters between 4 and 10 kHz. In male voices the region usually sits lower, in female voices higher. Because many languages have prominent 'sh' and 'ch' sounds, presets built around one speaker type can sit on the wrong band, so rather than trusting a generic 'vocal' preset you should find the band by ear for each speaker.
Check the Source First
Most sibilance problems come from mic placement and the recording chain: a voice speaking on-axis into a bright condenser has a problem before the mix begins. On production sound, a lavalier rubbing against fabric or sitting under clothing creates a different problem altogether. So before reaching for a de-esser, ask two questions: is this present throughout the speech, or only on specific words? If it's everywhere, a plugin makes sense; if it's a few words, manual treatment sounds cleaner.
Where Does the De-Esser Go?
Position in the signal chain changes the result a lot. My order on dialogue channels is: clip gain first, then EQ, then the de-esser, and the compressor last.
- ―After the EQ: if you've lifted the top end to add 'air', you've also lifted the sibilance. Putting the de-esser after the EQ keeps the cost of that brightness under control.
- ―Before the compressor: as the compressor lifts the low-level parts, it also pushes sibilance forward. If the de-esser comes first, the compressor works on an already balanced signal.
Finding and Setting the Band
- ―Engage the Listen/Monitor mode: most de-essers let you hear the detected band alone. Sweep it to where the 's' sounds sharpest.
- ―Choose between split-band and wideband: split-band reduces only the selected band and leaves the rest intact, while wideband turns down the entire signal. On dialogue, split-band is usually more natural.
- ―Set the threshold relative to the average level of the speech: only the harshest 's' sounds should be reduced by about 3–6 dB. If it works on every syllable, the threshold is too low.
- ―Limit the range: an over-suppressed 's' sounds like a lisp. The goal isn't to remove sibilance, but to bring it down to a level that doesn't bother the ear.
When the Plugin Isn't Enough: Clip Gain and Dynamic EQ
A de-esser treats all speech with the same sensitivity, yet sometimes only a word or two is a problem. In that case, isolating the sibilant in Pro Tools by splitting the clip and lowering that short section's clip gain by 3–5 dB is a much cleaner fix. It doesn't affect the rest of the clip and avoids the side effects a plugin can introduce across the dialogue.
Dynamic EQ sits between the two worlds: it applies a cut to a frequency band only when the threshold is exceeded. If the sibilant band is wide or the speaker changes, or if a de-esser makes the voice sound lispy, the softer curve of a dynamic EQ may work better.
Consistency Between ADR and Production Sound
ADR is recorded in a studio with a close mic, so it often carries more sibilance. If production and ADR clips sit side by side on the same dialogue line, you may need to de-ess ADR harder and production more lightly. Otherwise, at the cut, the brightness of the 's' sounds shifts and the ear reads it as a 'different room'.
A Short Checklist
- ―Is the problem across the whole speech, or only on specific words?
- ―Find the band with Listen mode; don't trust presets.
- ―Aim for 3–6 dB of reduction at most.
- ―Check with music and effects, at different listening levels.