Creature Vocal Design: From Human Voice to Monster
A good creature voice doesn't come from a library; it comes out of a human throat and gets processed in the right order. Recording, separating pitch from formant, layering, and keeping the performance alive — the practical steps.
When a short-film director asked me for the sound of 'the thing in the cave', the first thing I did was open the libraries. Twenty minutes later I gave up: the lion roar sounded like a lion, the bear growl like a bear, and none of them breathed with the creature's movement on screen. In the end I stepped in front of the mic myself, with a glass of water and a willingness to make ugly noises without embarrassment. Since then my rule is simple: a creature voice is a performance, not a processing chain. This article is about how I record that performance and turn a human voice into something non-human without killing it.
Character First, Sound Second
Before touching a plugin I answer three questions. Size: how big is it? That sets the fundamental pitch and the breathing rate; a large body means low resonance and slow respiration. Intelligence: is it near speech or purely instinctive? For intelligent creatures I keep syllable-like articulation and pauses; for animalistic ones a continuous, unbroken growl works. Physiology: does it have a mouth, teeth, is it wet or dry? A 'wet' creature lives on saliva and tongue sounds; a 'dry' one on glottal crackle and hissing air. These three answers make every decision that follows.
Recording: Capturing the Performance
I record with two mics: a large-diaphragm condenser at 20–30 cm for detail and air, and a dynamic (SM7B or similar) close to the mouth behind a pop filter for the body of the throat. The dynamic gives a dense low-frequency texture that doesn't fall apart when you pitch it down later. I aim for peaks between −18 and −12 dBFS; creature vocals are explosive and uncontrolled, and without headroom the best take is the one that clips. I record at 96 kHz when possible — not for the ears, but because once you drop the pitch an octave, content above 20 kHz becomes new harmonic material in the audible band.
- ―Record at least 10–15 variations of each type: growl, roar, breath, hiss, pain, death. In the edit, what you need is usually the take next to the one you thought was 'it'.
- ―Warm up before starting; after 40 minutes the voice is gone and there is no second session that day. Room-temperature water between takes, no coffee, no milk.
- ―Vocal fry and a throat growl are two different muscles. Fry is low and crackly, growl is noise-heavy; record them separately and layer them later.
- ―Record breaths as a separate pass with a separate energy. What makes a creature feel alive isn't the roar; it's the breath before and after it.
Pitch and Formant: Two Things That Must Be Separated
The classic beginner move is dropping the voice 12 semitones with varispeed or a plain pitch shifter. The result sounds like a 'slowed-down human': pitch falls, but so do the resonances of the mouth cavity (the formants), so the mouth turns into a giant cave and articulation blurs. The ear hears 'tape slowed down', not 'monster'. The fix is a tool that controls pitch and formant independently: Pitch II or X-Form in Pro Tools, or third-party options like Little AlterBoy, Krotos Dehumaniser, iZotope VocalSynth or Melodyne. My most common setting: pitch −5 to −9 semitones, formant only −2 to −4. The body gets bigger while the mouth stays human-sized, and the listener feels 'this thing is speaking'.
Layering: More Than Three Layers Is Mud
Processed human voice on its own usually feels thin. I think in roughly three layers. Bottom (sub): the same take pitched −12 to −19 semitones with everything above 120 Hz removed, or just the body of a lion/elephant roar — this layer is felt, not heard. Middle: the formant-preserved main performance; articulation, emotion and sync all come from here, and it sits loudest in the mix. Top (texture): saliva, teeth, cloth friction, hissing air, or a chicken/pig sound played fast — the detail that makes the listener feel it's organic. Each layer keeps its own envelope; instead of starting and ending everything together I delay the sub by 30–50 ms and start the texture a few ms early. That tiny offset mimics the physical reality of one sound coming out of one mouth.
Processing Order: Why It Matters
My chain usually runs like this: pitch/formant first (it works best on a clean signal), then saturation or light distortion (Decapitator, Saturn or a tape emulation — it thickens the harmonics), then dynamic EQ to soften the 'human' presence around 2–4 kHz, then a short convolution reverb or an IR (think of it as the creature's own ribcage: 40–80 ms, dry, resonant), and a compressor at the end. Putting pitch shift after distortion also shifts the harmonics the distortion created, and the sound drifts toward a synthetic robot tone — watch the order. Modulation effects (chorus, flanger) are usually absent; the moment a voice sounds 'plugin-y', the creature dies.
- ―Match the movement: a small pitch bend when the creature turns its head, a formant opening (+1) when the mouth widens. Automation is the cheapest realism tool you have.
- ―Never spread the sub layer into stereo; keep it centred and mono. Width belongs only to the texture layer and the reverb.
- ―If there's dialogue, reserve 1–3 kHz for the humans, not the creature. The creature can live below 200 Hz and above 5 kHz.
Breath and Silence
The most skipped layer is the creature during the moments when nothing happens. A continuous breath loop (3–4 different breaths of 4–6 seconds, crossfaded) keeps the creature present in the scene even when it isn't roaring. I process breath lighter than the roar: pitch −3, formant −2, a gentle high-pass, no sub layer. And cutting the breath right before the roar is the simplest tension trick there is; the listener doesn't notice the silence, but their body tenses. That's the essence of creature vocal design: the technique exists to make the performance invisible. If the listener hears the pitch shifter, we're in the wrong place; if they hear something breathing, the job is done.