Introduction
A polished vocal might pass through half a dozen plugins before the mix is finished, but the qualities that make it work are in no small part due to decisions made much earlier: in the performance, the room, the mic choice, and thousands of small choices, from EQ to compression settings, made by the engineer during tracking.
A vocal mix chain can only work with the material it receives. Compression can make a performance feel more present, EQ can heighten clarity, pitch correction can steady an uncertain note and reverb can place the singer inside a space, but each processor can only respond to what already exists in the recording. When the performance carries the right intent and the microphone captures it in the way the song demands and the engineer and vocalist intend, the mix becomes an exercise in revealing, controlling or building on those qualities. With a poor recording, processing increasingly becomes an attempt to manage problems, and ‘We’ll fix it in the mix’ should not be a mantra to live by!
The best vocal chains are built around decisions formed in reaction to what you’re hearing, not around a fixed sequence of fashionable plugins. A close, breathy pop vocal might need detailed level control and careful management of sibilance; a forceful rock performance might need fast, explosive compression, midrange emphasis and a short slapback delay; an intimate acoustic vocal might sound finished with a small amount of corrective EQ, some gentle compression and thoughtful use of a room or plate verb.
Waves has got multiple tools for each stage of that process, from noise reduction and pitch correction to compression, EQ, saturation, delay and reverb, and the aim of this guide is to show how those tools can fit together inside a coherent vocal workflow, beginning before the record button is pressed and ending with the final automation pass.
The vocal mix begins in the room
A good vocal recording starts with the singer feeling comfortable performing at their best.
That includes the obvious technical considerations, but also the headphone mix, lighting, room temperature, the amount of time available and, crucially, the way feedback is communicated between takes. A singer who’s straining to hear the track or their own voice, worrying about a lyric, or being repeatedly interrupted by an engineer fiddling with gear is not likely to produce their best performance, however good the mic happens to be.
The headphone mix needs special attention. The singer needs enough pitch and rhythmic information to stay connected to the arrangement, with their own voice coming through at a level that supports a confident delivery. Too much vocal in the headphones can discourage a vocalist from projecting, leading to an underpowered performance; too little of their own voice in the cans might encourage the opposite – pushing their voice, losing pitch control or varying their microphone distance. A bit of reverb in their cans can help some singers relax into the track, but watch that you don’t bake the verb into the performance – unless that’s what you’re going for.
Mic choice matters, but mic position can matter just as much. Moving a singer just a few centimetres can change the proximity effect, sibilance, room ambience and the balance between body and brightness more effectively than an EQ adjustment, a de-esser or an ambience-suppression plugin. A performer turning their head slightly away from the mic when delivering sharp consonants can soften an aggressive top end, while pulling back can give a powerful singer more room to move dynamically without overwhelming the capsule, but be aware that it may increase the room ambience in the recording.
There is no universal rule for how close to a mic a vocalist should be, because voices, mics and rooms all interact differently. Even with the same mic and room, a voice can change with the time of day, hydration, tiredness and other physical factors. A sensible starting point might be to place the mic around 10-12 inches from the vocalist, with a pop shield helping to stop the singer getting too close, but the final choice should be made while listening. A position that sounds great during quiet conversation might sound very different when the singer reaches the chorus.
Recording levels should leave comfortable headroom. There’s little value in driving every take as close to 0 dBFS as possible, and digitally clipped recordings can create problems that no conventional vocal chain can fully undo. Aim for a healthy signal whose loudest moments remain well clear of clipping, giving the singer freedom to perform and leaving the mix engineer room to work.
Record the performance you want to mix
Technical excellence is important, but you also need to ask whether the vocal take communicates the song, and this is a much more subtle detail to nail.
It’s easy to become distracted by technical imperfections during tracking, especially when the editing process later means that almost every syllable is individually available for inspection, but a slight change in tone, a breath that arrives early or a consonant that is less than precise can totally work in context when the performance has momentum and vibe. Watch for technical control causing emotional disconnection.
Comping is an essential part of the process when building a great take from top to tail, but it works best when you’re conscious of protecting the heart of the performance. Building every line from the most polished available word can produce accuracy, but at the expense of flattening the relationship between phrases, breaths and emotional intensity. Remember that a technically perfect take isn’t necessarily the best way to express the emotion contained in the song. Comping longer sections together can help to preserve a more convincing arc, using small repairs only where they genuinely improve the result.
The same logic applies to doubles and harmonies: their timing, pitch and consonants need enough agreement to support the lead without distracting the ear, but complete uniformity can remove the width and human movement that made them worth recording. Tight editing is useful when stacked vocals are clouding the rhythm or obscuring lyrics; looser alignment can sound larger and more natural when there’s space in the arrangement for it.
The goal of the recording session is to capture a vocal performance whose character is readily apparent and fits the vibe of the track.
Edit before building the plugin chain
Once the vocal takes have been selected and comped, a careful editing pass creates a stable foundation for the next steps.
Start by listening to the comp itself: check every edit at the boundary between phrases, listening for changes in room tone, headphone spill, microphone position and vocal timbre. Crossfades should be long enough to prevent clicks and abrupt changes, and short enough to preserve the beginnings of consonants and the natural timing of breaths.
Next, listen to the spaces between the lines. Complete silence can sound unnatural when you’ve otherwise got a consistent room tone, especially on exposed vocals. Cleaning obvious noises while keeping a believable background noise level usually produces a better result than cutting every gap to digital silence. Breaths can be reduced, reshaped or sometimes entirely removed, but they do form part of the phrasing, so taking away too many can make the singer sound disconnected from the performance. Exercise restraint!
Clip gain is really useful at this stage. A word or a single syllable that hits 10 dB hotter than the rest of the vocal line will force your compressor to react strongly – and that can change tone and ambience as a result. Manually pulling down peaks before the vocal hits the compressor means it can work within a narrower range of dynamics control, with more predictable results. Quiet syllables can be boosted for the same reason, especially when intelligibility is being lost at the ends of lines.
This doesn’t mean turning the waveform into a solid block. The vocal should retain its musical dynamics. The point is to address isolated level jumps (up or down) that could destabilise processors further down the chain.
When vocal cleanup plugins help
A recording made in an acoustically treated, well-designed studio might need only a conventional comp and edit. Home-recorded vocals or location performances can need more work, with computer fans, traffic, air conditioning, room reflections or broadband background noise finding their way into the recording. This is where tools like Waves Clarity Vx can be really useful.
Clarity Vx is designed to separate voice from background noise, whilst Clarity Vx DeReverb helps to reduce unwanted room sound or recorded reverb. Used on the right material and the right set of problems, both plugins can do a really good job of recovering recordings that would otherwise be difficult to improve through more conventional means, like EQ or gating.
But have a clear target in mind when you use these types of processors. A small reduction in noise, accepting that you can’t completely eliminate it, can be more viable than aggressively cleaning a track and leaving mangled consonants, squelchy breaths and high-frequency detail loss in your wake. There’s always a trade-off, so listen to the vocal in the context of the full track as well as in solo, because low-level room noise that seems distracting in solo might completely disappear once the music is playing.
Where noise is audible only during gaps, careful editing and gentle expansion might serve you better and preserve voice details more effectively. You’ll need to judge for yourself on a case-by-case basis.
Pitch correction: support the performance
Pitch correction isn’t only a technical corrective tool – it can be transparent, but you might want to use it audibly and get creative. There are no absolutes when it comes to the right or wrong way to use it, but bear in mind the context of the production you’re working on and how it’s going to fit aesthetically.
Check your key signature and listen through the whole performance. Some ‘pitch problems’ are actually timing problems in disguise, unstable levels, unusual melodic movements or even deliberate transitions between notes, and the last thing you want is your pitch correction plugin ‘correcting’ something that should be preserved. An experienced and accomplished singer might intentionally approach a note from below, allow vibrato to spread around its centre or briefly pass through a non-scale pitch as part of the phrase, and correction that treats every departure from a predefined allowable set of notes as an error can make a musically expressive performance feel artificially constrained.
Plugins like Waves Tune Real-Time provide automatic pitch correction with adjustable note transitions, speed, tolerance, formant handling and vibrato control, and Tune Real-Time’s low-latency design means it can be used during vocal tracking or live performances too, not just mixing.
For natural correction, start more gently than you think you need. Make sure you’re set to the right key and scale, disallow any notes you feel don’t fit the melody, and then adjust the speed and transition of the correction until unstable notes start to settle without every movement snapping towards the note centre. Slow songs, exposed vocals and performances with wide and noticeable vibrato often need careful transition time settings to keep their shape.
A modern hard-tuned effect calls for a different approach altogether! Fast correction, tightly controlled notes and deliberate wrangling of formants can become part of the vocal’s identity, especially when the singer hears the processing whilst performing, in which case the plugin affects phrasing as well as pitch. Vocal slides, held notes and transitions can be performed with the response of the tuner shaping the vocalist’s delivery.
Real-time correction is fast and definitely has its uses, but automatic processing can’t know which moments matter most. Listen out for notes pulled in the wrong direction, transitions that have become abrupt, and consonants or breaths that briefly trigger unstable or inaccurate detection. You might need some automation or detailed editing to avoid these isolated problems while leaving the rest of the performance relatively untouched, or a different kind of pitch correction plugin like Melodyne, Brainworx’s Crispy Tuner or Waves Tune.
Establish the vocal’s working level
The next step, before reaching for EQ and compression, is to balance the vocal with the arrangement.
The first balance is provisional – you’re not aiming for perfection – and it reveals what the track is asking of the vocal. A voice that sounds thin when soloed might sit perfectly above dense guitars; a warm, full vocal might seem congested or muddy beside a rich-sounding piano, synth pads and layered backing vocals. Processing vocal tracks on their own, in solo, is rarely the way to go if you want them to ‘sit’ naturally in the mix.
With your full arrangement playing, bring up the vocal fader until the words, tone and emotional emphasis are clear, then move around the song. Notice where the vocal disappears, where it feels too close and where the arrangement fights for the same space. This is a much better way to figure out what needs to change to get it sitting right.
Some engineers like to place a trim or gain plugin at the beginning of the chain, so they can control the level feeding the compressors and other input-sensitive processors that follow. Others prefer to use clip gain, sometimes working right down to individual words or syllables, so the vocal reaches the plugin chain with a more consistent level. This can be especially useful when dynamics presets are involved, because compressors, limiters and gates respond very differently depending on how hard they are being driven and how uneven the incoming signal is.
Corrective EQ: clear the path
Vocal EQ is most commonly used for a couple of specific reasons in a mix: to control distracting features in the recording, like resonances picked up from room excitation or produced by the vocalist, and to shape the frequency content of the voice to help it sit in the right place in the arrangement.
A high-pass filter can remove rumble, mic stand movement and low-frequency energy that doesn’t contribute much to the vocal, but the cutoff frequency should be chosen carefully by ear. Don’t automatically roll off by default – processing should always be intentional, with a specific goal in mind. Setting an HPF too high can weaken the body of a voice, especially on lower male vocals or intimate performances where mic proximity is a desirable part of the sound. Roll away only what you don’t need, and err on the side of caution!
Low-mid buildup can happen anywhere between about 150 Hz and 500 Hz, but the precise frequency varies enormously with the singer, mic and room. Listen carefully to identify the character you’re hearing, then make the smallest EQ move that creates a meaningful improvement. A cut with a wide Q might help with overall ‘boxiness’; a narrower cut can help with specific resonances, although very deep notches can call attention to themselves when the vocalist moves between notes, so keep listening and exercise restraint!
The upper mids need similar care. Energy at around 2 to 5 kHz helps with intelligibility and brings the vocal forward, but too much can sound edgy and harsh, making consonants brittle, increasing listening fatigue and causing the vocal to fight with guitars, snare or other instruments in similar areas of the frequency spectrum. Boosting high frequencies with a shelving filter can add openness and air, but it’ll also boost sibilance and mouth noises, and can increase room ambience – so careful listening to what’s going on around the performance is a must.
Dynamic EQ can be useful when the problem changes with the performance. Waves F6 gives you six independently configurable bands that can operate conventionally or as a dynamic EQ, with attack and release controls on each band, meaning a harsh frequency can remain untouched where it’s not a problem and be reduced when it becomes excessive.
That’s useful for a vocal that gets too strident on big notes, develops a low-mid bloom on certain vowels or needs some presence control. When a deep static cut seems necessary, dynamic EQ can sometimes give more natural results – leaving much of the performance untouched and engaging only where it’s needed.
Another approach is to use a channel strip plugin, like Waves’ SSL E-Channel, which combines filters, EQ, dynamics and gating in one familiar console-style workflow. Its midrange bands can come in really handy when the vocal needs broad tonal shaping and a decisive sense of direction, while the integrated compressor can provide some initial dynamics control and be paired with another compressor or two further down the chain.
Compression: control, movement and character
Vocal compression is undeniably about dynamic range, but it’s also about so much more than that: it’s about sustain and body, presence and attack, saturation and colour and the way each phrase moves through the track.
A fast attack can grab sudden volume spikes and create density, but set it too fast and consonants can be overly softened, while the small transients that help articulation can suffer. Slower attack times allow more of the front of each word through, increasing presence and apparent energy at the expense of some dynamic control. Release times determine how quickly the compressor ‘lets go’ of the signal: too slow, and the vocal might stay ducked for too long after a loud phrase; too fast, and the level can move too audibly, or ‘pump’, between syllables.
The sweet spot for attack and release settings depends on the recording.
Using the CLA-76 for peak control and urgency
The Waves CLA-76 models two 1176-style compressors, the more vibrant Bluey and the somewhat smoother Blacky, with extremely fast attack and release ranges, an input-driven fixed compression threshold, fixed 4:1, 8:1, 12:1 and 20:1 ratios, a parallel Mix control, and a Comp Off mode that runs the signal through the ‘hardware’ circuitry whilst bypassing gain reduction to retain the analogue-modelled colour.
The CLA-76 really comes into its own on vocals, catching fast peaks and delivering an up-front sound with added ‘presence’ and urgency. A great starting point is a 4:1 ratio, a moderately slow attack and a moderately fast release, then you can dial in the exact sound you’re after. Remember that the attack and release numbering runs from slower at 1 to faster at 7. If you’re still a little new to compression, make it work hard with lots of gain reduction so you can really hear the effect, then back off once you’ve set the attack and release. In fact, with an 1176 – and plenty of other compressors – don’t get hung up on the gain reduction needle; it’s much better to listen carefully and tune in to how the compressor grabs the fronts of words and lets go as the release characteristics reveal themselves.
CLA’s ‘Bluey’ can bring a vocal forward with a little more edge and midrange attitude than the ‘Blacky’, which sometimes suits material that benefits from firm control with a little less character. These are tendencies, mind you, not absolutes; the best choice is the one that fits the material you’re working with. Whilst you’re learning the tools, experiment – and trust your instincts.
Using RVox (Renaissance Vox) for fast, focused levelling
RVox is a very different kind of plugin. Under the hood it’s a relatively slow-attack, fast-release compressor, and in that sense it’s not dissimilar to the CLA-76, but that’s where the similarity ends. It has just three main controls: two for the compressor – threshold and makeup gain – and one for the gate. It’s designed specifically for vocals and dialogue, and its simplicity makes it really quick to use. Its non-adjustable attack and release settings also make it a great fit for vocals in many, many cases.
R-Vox can add density and hold a vocal steady in the mix with very little adjustment. Pull the compression control down until things get more consistent, and check that low-level details haven’t been brought forward excessively. Sometimes you want low-level details – breaths, mouth noise, headphone spill and room ambience – to come up, because that can give extra energy to a performance, but listen in context and be aware: the simplicity of the controls doesn’t remove the need to listen!
It also works well as the second stage in a serial compression chain: a CLA-76 to catch the sharpest peaks and add character, then R-Vox for some extra levelling. By the way, RVox also has a limiter built into the output stage, so no matter how hard you drive it, you won’t get digital clipping at the output. Whether you hear distortion from the limiting is another matter, so don’t drive it hard without listening carefully!
MV2 and low-level detail
Waves MV2 is both a traditional ‘downward’ compressor and an ‘upward’ compressor that raises the volume of quieter material falling below a user-defined threshold. On a vocal, this is a brilliant tool; it can improve audibility at the ends of phrases, where vocalists can dip in volume, and help to create a dense, finished sound with only a couple of key controls. WPR’s MV2 review explores its upward and downward compression in more detail.
The Low Level control is a lot of fun, and a powerful thing! It can reveal diction and intimacy, breaths and room reflections, as well as unwanted background noise and poor edits. A clean, controlled recording might respond beautifully; a noisy one, not so much. But it can add a huge amount of character and vibe by emphasising the physicality of a performance, and sometimes bringing up less desirable artefacts is a trade worth making for more vibe.
De-essing in context
Sibilance is partly a recording issue, partly a performance characteristic, and partly a consequence of processing.
Bright microphones, close miking and direct alignment with the capsule can exaggerate “s”, “sh” and “ch” sounds before the signal reaches the DAW, and high-frequency EQ and compression can make those consonants even more pronounced. This is why de-essing, like so many other processes, is best done in the context of the mix – not in solo – and why it’s worth experimenting with where the plugin sits in the vocal chain. There’s no definitive right or wrong signal-flow order when it comes to de-essing.
One of Waves’ more useful tools for dealing with excess sibilance is, surprise, surprise, Waves Sibilance. It uses something Waves calls ‘Organic ReSynthesis technology’ to identify and reduce sibilance whilst keeping more of the original vocal’s surrounding signal. I have several de-essers I regularly turn to because each voice is different, and one that works brilliantly on one recording might fall short on another. Waves Sibilance is often the one I choose.
You’ve got a few key controls at your disposal here: ‘threshold’ and ‘range’ determine the amount of gain reduction, whilst ‘detection’ and ‘mode’ help you process the right part of the frequency spectrum. Listen carefully to the essiness that’s causing the problem – is it a thin, specific frequency, or a wider ‘sh’ sound? Use the detection and mode controls to fine-tune what you’re reducing.
The goal is not the complete disappearance of every “s”! Too much reduction can make diction sound soft, lispy or disconnected, so tread carefully and keep checking how much de-essing you’re doing by switching the plugin in and out of bypass. It’s easy to lose perspective with any EQ-based tool, and de-essing is essentially frequency-dependent compression or dynamic EQ.
Think about where you want the de-esser in your chain: placing it before compression can prevent sharp consonants from driving the compressor too hard, whilst placing it afterwards can control sibilance that compression and tonal shaping have brought forward. Just as some vocals benefit from two light stages of broadband compression, some benefit from two stages of de-essing in different positions: one early in the chain to control the strongest esses, and another later to manage the overall brightness. Some mix-bus processors even include de-essing stages – Plugin Alliance’s bx_masterdesk, for instance. Really, it’s whatever works.
Dynamic EQ can also work well as a de-esser in certain situations. As with a dedicated de-essing plugin, the result depends on the voice, the recording and the specific problem you are trying to control, but Waves’ F6 is well suited to this kind of work because each band has adjustable attack and release. That means you can shape how quickly it responds to sibilance, how naturally it lets go, and whether it feels like transparent control or a more obvious tonal change.
Saturation and tonal character
Once the vocal is dynamically where you want it and sitting broadly in the right place in the mix, you might want to experiment with saturation to give it a little more density, texture or character.
There are thousands of saturation plugins on the market, and saturation comes in many varieties, but one solid place to start is Waves’ Abbey Road J37 tape-emulation plugin. It models the valve circuitry, magnetic tape formulations and tape behaviour of Abbey Road’s Studer J37 machines, and gives you control over tape speed, saturation, bias, wow, flutter and noise, as well as a tape-delay section.
It’s capable of producing some very obvious effects, but on a lead vocal I’ll often use it much more subtly. ‘Saturation’ doesn’t always need to equate to riotous distortion, and used carefully the J37 can add a bit of harmonic activity, soften the edges of the recording and help the midrange feel slightly more connected.
Try raising the input until you begin to hear the centre of the voice filling out, compensate for the increased level at the output, then switch the plugin in and out of bypass to check that you really prefer what it’s doing. Louder nearly always sounds more impressive at first, and it’s easy to mistake an increase in volume for an improvement in tone, so level-match the processed and bypassed signals before deciding which you prefer.
You can, of course, push it much harder. More input drive, a different tape speed, increased wow and flutter and a short tape delay can turn the J37 into an audible ‘production’ effect. I’m more likely to do this on doubles, ad-libs, backing vocals, telephone-style passages, a short section where I deliberately want the perspective to change, or across a vocal where a more ‘opinionated’ mix is called for. It can also distinguish a response vocal from the lead without relying purely on EQ.
The modelled tape noise is an even more subjective matter of taste, and whether it adds anything useful depends on what you’re trying to achieve. On a deliberately vintage production it might contribute to the illusion, but across a large number of tracks the noise can build up quickly and become too much! Most of the time, the saturation and dynamic behaviour give me the character I’m after, and I’ll leave the noise off – but it’s there if you want it.
Vocal Rider and the final level pass
Compression can control the ‘internal’ dynamic range of a vocal and change the way individual words and phrases move, but it won’t necessarily put every line or syllable at the right level against the arrangement. Remember, compression isn’t just a levelling tool, it’s a vibe tool too, and there are sometimes better ways to place a vocal against a track whose level is constantly shifting.
A chorus might need the vocal to come up because the guitars, drums and backing vocals have expanded around it. A quiet verse might need the opposite. Sometimes one word needs to be pushed up because it carries the meaning of the line, even though it is already perfectly audible. Those are musical and balance decisions rather than compression problems, and this is where vocal riding and volume automation come in.
Imagine a finger moving your volume fader up and down to hold the vocal level steady in the mix – that’s Waves Vocal Rider. It automatically adjusts the vocal within a range you define, pulling down louder sections and pushing up quieter ones. It can also listen to the instrumental mix through its sidechain input, allowing it to respond to the changing level of the music as well as to the vocal. But a processor making a technical decision isn’t the same as a mixer making the right musical one, so I tend to give Vocal Rider a fairly controlled operating range. It can create a useful first pass, but watch what it does with breaths, sustained notes, quiet line endings and gaps between phrases, and make sure it isn’t chasing details you would rather leave alone or making decisions that don’t work with the intention of the line or the vocal balance.
Vocal Rider can’t know that the final word of a line is meant to fall away, or that a lyric needs to lean towards the listener before the chorus arrives, so think of it as a way of reducing the corrective riding needed rather than a complete automation-replacement tool. Give the vocal a pass of Vocal Rider, then write volume automation over what it’s done, fine-tuning section changes and word-by-word or syllable-by-syllable boosts and cuts yourself.
Another way of levelling and adding emphasis to vocal tracks is to use clip gain and conventional volume automation from the outset. I use both approaches, depending on the vocal, the time available and how complicated the arrangement is.
Whichever way you do it, though, don’t underestimate this stage. A carefully ridden vocal can sound more finished, more emotionally connected and more expensive without another tone-shaping plugin being added to the chain.
Reverb and delay: creating a space around the voice
A completely dry vocal can feel as though it’s sitting on top of the track rather than inside it, but too much reverb can push it back too far, soften the diction and reduce the intimacy you worked hard to capture. As ever, there’s a balance to find.
I’ll normally put vocal reverbs on aux returns rather than directly on the vocal insert. This lets the ‘dry’ and ‘wet’ signals be balanced independently, allows several tracks to share the same space, and gives the reverb return its own EQ, compression and automation. EQ is pretty essential to shaping the colour of the reverb, and automation is a really useful creative tool for shaping the returns.
Pre-delay is one of the most useful controls for keeping a vocal clear. By setting a small gap between the dry voice and the beginning of the reverb, you can keep the front of the word up close and intelligible whilst also having the ambience develop behind it. A longer pre-delay can create a separated, polished sound (common in pop productions), whilst shorter settings tend to place the singer more firmly inside the acoustic space.
There isn’t a correct pre-delay value for vocals. The song tempo, vocal phrasing, reverb length, density of the arrangement and speed of the lyrics all affect the result, so set it while the track is playing and listen to whether the reverb is supporting or smearing the performance.
Filtering the return can make an enormous difference. Low-frequency reverb can take up a lot of space without contributing much to the mix, whilst a very bright tail can exaggerate sibilance, draw attention to mouth noise, compete with cymbals or percussion, and sometimes reveal weaknesses in a lesser reverb algorithm. Darker reverbs can be turned up further before they interfere with the words, which can make the perceived space feel larger. Having the reverb on its own aux gives you the chance to EQ and shape it with all of these dimensions in mind.
Waves H-Reverb is a flexible verb capable of emulating a number of spaces and reverb types, including short rooms, plates, larger halls and more obviously designed effects, with a good level of control over the envelope, modulation, filtering and dynamics. That adjustability is useful, but don’t feel you need to use every control. Start with the length, pre-delay and tonal balance of the return, because those three decisions will usually get you most of the way towards the space you need.
Delay is another way of creating depth and width, often with less continuous wash than reverb, leaving the mix feeling more spacious. Waves H-Delay includes tempo-synced and millisecond delay times, feedback, filtering, modulation, ping-pong operation and several analogue-style modes.
A short mono slapback delay can add dimension and energy without making the vocal feel obviously wet, and it can be particularly effective on rock, indie and retro-feeling productions. Longer delays can support the rhythm of the track, but they need to be timed and filtered carefully so that the repeats don’t obscure the lyrics that follow.
Delay and reverb automation shouldn’t be overlooked – it’s an area that really earns its keep. Instead of leaving the same delay running throughout a song, try sending the last word of a line into it, increasing the feedback for a transition, or allowing the repeat to fill a gap in the arrangement. The same goes for reverb: the verse might need to feel close and contained, whilst the chorus might benefit from a broader return, or one exposed lyric might become more powerful if the ambience disappears altogether. Use automation to draw attention to special moments and support the emotional flow of the track.
By the way, the J37 that we looked at for saturation also has a really effective, simple-to-use tape-delay section. It can be especially good for short, coloured repeats and vintage-style slapbacks. H-Delay is often my go-to for precise rhythmic control or more modern stereo delays – it’s actually really versatile.
We’ve got detailed WPR reviews of the J37 and H-Delay if you’d like to look more closely at the way their individual controls work.
Three real-world Waves vocal chains
These chains are intended as starting points. The order makes sense for the particular job I’ve assigned to each processor, but another recording might lead me to move things around, leave half of them out or use a completely different tool. The processing choices you make must always be in service of the material you’re working with.
1. An intimate, natural vocal
For an exposed singer-songwriter performance, acoustic production or restrained ballad, I would be wary of processing away the small movements that make the vocal feel personal.
Editing and clip gain → gentle pitch correction where necessary → F6 corrective EQ → light R-Vox compression → Sibilance → Vocal Rider or manual automation → short H-Reverb room or plate (on an aux) → occasional H-Delay throw (on an aux)
Begin with the edit and clip gain, making sure that the phrasing and level changes still feel like a cohesive performance without taking all the life out of it.
Where pitch correction is needed, use it gently enough to keep the singer’s transitions and vibrato. F6 can then control any low-mid bloom, upper-mid hardness or occasional resonances without stamping a new tonal personality across the vocal.
R-Vox is useful here because a small amount of reduction can help to hold the recording steady without requiring a complicated setup. Keep checking that you haven’t made the vocal so consistent that it stops breathing with the song!
I would normally set the de-esser after the main tonal and dynamic processing in this chain, because those processes may have changed the apparent brightness and brought the sibilance forward. A short room or plate can then provide enough acoustic connection to stop the vocal feeling isolated, with delay appearing only where a phrase genuinely benefits from it.
2. A controlled modern pop vocal
A dense pop arrangement can ask a lot from a lead vocal. It might need to sound bright, intimate and apparently effortless whilst remaining stable against synths, drums, guitars, backing vocals and production effects that are constantly changing around it.
Clarity Vx where required → Tune Real-Time → clip gain → F6 → CLA-76 → R-Vox → Sibilance → J37 → Vocal Rider → H-Delay and H-Reverb sends
If the recording contains background noise or distracting room sound, deal with that near the beginning. Cleanup processors are usually better placed before compressors, saturation and high-frequency EQ have had a chance to emphasise the unwanted material.
Tune Real-Time can then handle the pitch correction. That might be subtle stabilisation, or it might be an audible effect; the important thing is to decide intentionally and lean into your choice.
Use clip gain and F6 to create a reasonably controlled signal before compression. The CLA-76 can catch the faster peaks and add some urgency, and R-Vox can provide a second, more straightforward stage of levelling. Don’t assume both compressors need to work hard: let your ears guide your decisions. Two compressors each doing a little can sound more natural than one doing everything – if that’s what you’re after.
Set Sibilance after those stages if the compression has brought consonants forward, then use the J37 for as much (or as little) harmonic density as the production needs.
Vocal Rider or manual automation can take care of the changing relationship between the vocal and the arrangement. I’d usually automate the effects by section as well: maybe a tighter, drier verse, a broader chorus, a few delay throws and a longer reverb tail leading into the final section.
3. A forward rock vocal
A rock vocal often needs to compete with dense guitars and energetic drums without becoming thin, harsh or excessively loud. It can benefit from the extra urgency, presence and bite that a ‘76-style’ compressor like the CLA-76 can bring.
SSL E-Channel → CLA-76 Bluey → Sibilance → R-Vox or MV2 where required → J37 → manual automation → short H-Delay slapback send → compact H-Reverb plate or room sends
The SSL E-Channel, which naturally sculpts the sound in subtle ways even with the EQ flat, can set the general tone. Use it to remove low-frequency energy that isn’t making a useful contribution, shape the midrange so that the vocal has somewhere to live amongst the guitars, and make the broad decisions that place it where you want in the frequency landscape. Other tools, like FabFilter Pro-Q or Waves F6, are better suited to narrow notches.
The CLA-76 Bluey brings fast dynamics control and forward energy, but pay careful attention to the attack. Too fast can take bite and intelligibility away from the fronts of words; too slow can let too much of the initial peak of hard syllables through the compression circuit, meaning you’re not getting the benefit. There’s a sweet spot where the vocal attack can be heightened and the compression pulls the rest of the phrase forward. Don’t be afraid to fiddle with the attack and release settings whilst applying far too much gain reduction, so you can really hear what you’re doing.
De-ess after the CLA-76, then decide whether you actually need another levelling stage. R-Vox can hold the vocal more firmly, while MV2 is useful when the quieter details need lifting and the recording is clean enough to withstand it. Neither has to be present just because it appears in the chain.
J37 can add thickness and help the vocal feel less separate from the guitars. A short slapback will often provide more useful size than a long reverb, and a compact room or plate can stop the voice feeling ‘pasted’ onto the front of the mix.
Lastly, ride the vocal. Dense sections might need important lines lifted out, while gaps in the arrangement might mean you can ease it back a bit. There’s no single fixed compressor setting that can make all of those decisions for you.
How many vocal plugins do you actually need?
Honestly, you don’t need very many.
Most DAWs include perfectly usable EQs, compressors, pitch editing tools, delays and reverbs, and unless another plugin gives you a particular sound, workflow or degree of control that is genuinely useful, chances are you could mix just fine with the DAW’s own stock plugins.
A third-party plugin earns its place when it solves a problem better than the tools you already own, gets you results more quickly, or produces a character you value and can’t find another way. It doesn’t earn its place simply because somebody included it in a screenshot of their vocal chain 🙂
For a relatively compact Waves vocal toolkit, I’d probably start with the following:
- A flexible EQ: F6 can handle conventional EQ, dynamic control and some de-essing duties.
- A principal compressor: R-Vox is exceptionally quick to set up and use when you want density and straightforward levelling. The CLA-76 gives you more control over the movement and character of the compression. If I could have only one – I’d pick R-Vox. It’s so quick to use, and in 99% of cases it works brilliantly.
- A de-esser: Sibilance is a strong dedicated option, although F6 may already cover the job on some voices.
- A pitch processor: Tune Real-Time is useful for automatic correction, monitoring through pitch correction and more audible tuned effects.
- A delay: H-Delay is really versatile and sounds great. It can cover slapback, rhythmic repeats, filtered throws and plenty of more obvious production effects.
- A reverb: H-Reverb gives you a great deal of control, but there’s no need to buy it purely for completeness if the reverb supplied with your DAW already creates the spaces you need.
Clarity Vx, Vocal Rider, J37 and MV2 are worthwhile additions if their particular jobs crop up often enough in your work. A home-recording engineer dealing with inconsistent rooms may value Clarity Vx far more than someone routinely recording in treated studios; likewise, an engineer who enjoys automating the volume of every word by hand probably won’t benefit from Vocal Rider in the way that someone working quickly through a large number of mixes will.
Use the smallest chain that gives you the results you’re after. Keeping it small and manageable makes it easier to understand, tweak and troubleshoot when something doesn’t sound right.
Final thoughts
A vocal mix doesn’t begin when you insert the first plugin.
It begins with the singer, the room, the microphone, the headphone balance and the decisions made while the performance is still being recorded. Editing, pitch correction, EQ, compression, de-essing, saturation, effects and automation are all continuations of that same process, and every stage changes what the next one has to work with.
A strong performance, captured in a suitable space with a microphone that complements the voice, might need surprisingly little processing. A recording made in difficult circumstances can still be improved enormously through careful editing, cleanup and dynamics control, but the plugins will inevitably be working harder and the compromises will be more audible.
Waves has got tools that cover every part of the journey. Tune Real-Time can gently steady a performance or become an obvious part of its sound. F6 can control tonal problems that appear only on certain notes. CLA-76, R-Vox and MV2 each approach dynamics from a different direction. Sibilance deals with consonants, J37 supplies colour and delay, Vocal Rider helps with level consistency, and H-Delay and H-Reverb create movement and space.
That doesn’t mean they all belong on every vocal.
Start with the performance you’ve actually got, listen to it against the arrangement and decide what’s preventing it from doing its job, then deal with that first. Preserve the character that already works, and add another processor only when you can explain or hear in your head what you want it to contribute.
The plugin chain should be the result of those decisions, not the starting point for them.













