The Vocal Chain

Microphone

Everything starts here, and no amount of processing downstream fixes a bad source. The microphone isn’t just capturing a voice — it’s making an interpretive decision about it. A large-diaphragm condenser like a Neumann U87 or an AKG C414 tends to flatter most vocal timbres, emphasising presence and air in a way that translates well to recorded music. A dynamic microphone — an SM7B, say — sits closer to the voice in a different sense: it rolls off some of that sheen and handles high SPLs without flinching, which is why it became the go-to for broadcasters and, famously, for Michael Jackson recording Thriller.

The right choice depends on the voice, the room, and the genre. A bright, airy condenser in an untreated room will pick up every flutter and reflection. A darker dynamic in the same room might give you something usable. Know what you’re capturing before you decide how to shape it.

Preamp

The microphone signal coming out of a condenser is tiny — too small to record cleanly without amplification. The preamp brings it up to line level, and in doing so, it adds character. This is often understated. A clean, transparent preamp like an SSL or a Neve 1073 clone will colour the sound differently. The SSL tends toward a kind of clinical precision; the Neve adds warmth and a gentle harmonic density that many engineers describe as making vocals sound more “expensive.”

If you’re working in the box entirely, your audio interface has a preamp built in. Entry-level interfaces have come a long way — a Focusrite Scarlett or an Apollo Solo will give you a clean signal that serves most purposes. The preamp stage is where gain structure matters: too little gain and you compensate with noisy amplification later; too much and you clip before the signal even hits your DAW.

Compression

Vocals are dynamic by nature. A singer will whisper a verse and belt a chorus, and the difference in level can be enormous — far more than a final mix can accommodate. Compression manages that range, bringing up the quiet moments and reining in the loud ones, so the performance sits consistently in the track without constantly fighting for its place.

The LA-2A — a levelling amplifier originally designed for broadcast — has become a near-universal first choice for vocal compression because its programme-dependent attack and release times respond to the music rather than fighting it. The 1176 offers more aggressive, faster control and a particular harmonic saturation when pushed, which many producers use deliberately on rock and hip-hop vocals. The order here matters: compression before EQ means you’re shaping a signal that has already been dynamically controlled. That’s usually what you want — EQ-ing before compression can cause the compressor to react differently to boosted frequencies, which can introduce unpredictability.

That said, some engineers run a second EQ before compression intentionally to manipulate how the compressor responds — cutting a problematic low-mid, for instance, so the compressor isn’t triggered by it. These aren’t rules, they’re tools. Know why you’re breaking the chain before you do.

Equalisation

Once the dynamic range is under control, EQ sculpts the tonal character of the voice. A high-pass filter — typically set somewhere between 80Hz and 150Hz depending on the singer — cuts the low-frequency content that doesn’t belong to the vocal: room rumble, mic stand vibration, the low end of a chest-heavy baritone that might muddy the mix. This alone can make a vocal sit more clearly in a track without touching anything above.

Above that, EQ becomes a conversation between the voice and the arrangement. A boost around 3–5kHz adds presence and intelligibility — the consonants cut through. Somewhere between 200Hz and 400Hz is where many vocals carry an unwanted thickness, a kind of boxiness that can be gently reduced. Boosting above 12kHz adds air, though this can expose noise in the recording if the source isn’t clean.

The reason EQ typically comes after compression in the vocal chain is that compressors are frequency-sensitive. Boost a lot at 3kHz before compression and the compressor may over-react to those frequencies, clamping down harder than you intend. Post-compression EQ is shaping a more stable, predictable signal.

De-essing

Sibilance — the hiss of “s,” “sh,” and “t” sounds — is a physiological fact of speech that microphones, particularly condensers with a presence peak, can exaggerate brutally. A de-esser is a frequency-specific compressor that targets a narrow band, typically between 5kHz and 10kHz, and compresses only when energy in that range exceeds a threshold. Done well, you don’t hear it working. Done badly, vocals develop a lisp or a lispy, pumping quality that distracts from the performance.

De-essing sits after EQ for a reason. If you’ve boosted presence frequencies to add clarity, you’ve likely also boosted the region where sibilance lives. Addressing that after the fact — with a de-esser dialled to exactly the frequency that’s causing the problem — is more surgical than trying to manage it with broad EQ moves that would sacrifice the clarity you just worked to add.

Reverb and Delay

A dry vocal is almost always wrong for a finished mix. Without some form of spatial treatment, the voice sits in a different acoustic space than every other element in the production — instruments recorded in studios, samples made in different rooms, synthesisers that exist in no room at all. Reverb and delay don’t just add “wetness”; they place the vocal in a world.

Delay tends to sit the vocal forward — it adds depth without washing out definition. A short slap-back delay (think early rock and roll, Sun Records, Elvis Presley in the mid-1950s) thickens a vocal without pushing it back in the mix. A longer rhythmic delay, synced to the tempo of the track, extends phrases in a way that can feel compositional rather than merely spatial. Reverb does something different: it suggests the size and character of a room. A short plate reverb keeps things intimate and professional. A long hall reverb can make a voice sound enormous, which is either the point or a problem depending on the music.

These effects come last in the chain — or more typically, they live on a separate send/return bus — because they’re designed to process a signal that has already been cleaned, compressed, and shaped. Sending an unprocessed vocal into reverb means the reverb tail carries all the problems: the dynamic inconsistencies, the sibilance, the low-end rumble. Process first, then place the result in space.

There’s also a practical argument for using sends rather than inserting reverb directly on the vocal channel. A send allows multiple instruments to share the same reverb space, which is how acoustic environments actually work — everything in a room sounds like it’s in the same room. Vocals that share reverb and delay with other elements feel like they belong to the same recording. Vocals on a private reverb often sound like they were added later, because sonically, they were.

This chain — microphone, preamp, compression, EQ, de-esser, reverb and delay — is not a formula. It’s a framework built from the accumulated decisions of thousands of sessions across decades of recorded music. Every stage has its logic, and the order has a reason behind it. Once you understand why the chain is sequenced this way, you’ll know exactly when to break it.

Explore how these stages interact in a full mix context with the Resonillator Mixing Guide.

Resonillator is free and always will be. If this was useful, consider supporting on Patreon or Buy Me a Coffee.

← All Transmissions