How to Mix Spoken-Word Vocals

Web Admin Avatar

·

7 min read

How to Mix Spoken-Word Vocals

To mix spoken-word vocals well, the goal is clarity and consistency, not the polish you chase on a sung lead. Work in this order: clean up the recording, carve problem frequencies with EQ, compress for an even level, tame sibilance, add a little presence, then set a loudness target for the platform. That chain keeps every word intelligible from a phone speaker to a good pair of headphones.

Violet Recording is reader-supported — we may earn a commission from links on this page, at no extra cost to you.

Spoken word — podcasts, voiceover, narration, audiobooks — lives or dies on being understood. There’s no melody to hide behind and no beat to fill the gaps, so noise, boominess and harsh “s” sounds stand out far more than they would in a music mix. The good news is that the moves are simple and mostly repeatable, and your DAW’s stock plugins can do almost all of them. Here’s how to approach each stage.

Clean up the recording before you mix spoken-word vocals

Fixing problems at the source is always cheaper than processing around them, so start with a clean signal. First, roll off the sub-lows with a high-pass filter. Speech has almost no useful energy below the fundamental of the voice, so a high-pass somewhere around 70–100 Hz for a deeper male voice, or 90–120 Hz for a higher voice, removes rumble, handling noise and air-conditioning hum without thinning the tone. Set it by ear — sweep it up until the voice starts to sound small, then back off.

Next, deal with plosives and breaths. Hard “p” and “b” sounds create a low-frequency thump; your high-pass helps, but for stubborn pops you can automate a quick volume dip or use a dedicated de-plosive move on that word. Loud breaths can be pulled down by 3–6 dB with clip gain or gentle automation — don’t remove them entirely, or the delivery sounds robotic.

Background noise is the one job where stock tools sometimes fall short. A light broadband noise reduction of a few dB is usually invisible; push it hard and you get a watery, underwater artefact. If you’re rescuing genuinely noisy field recordings, a specialised restoration plugin will pull cleaner results than most built-in de-noisers — this is one of the few spoken-word tasks where a third-party tool genuinely earns its place. (Plugin Boutique carries the well-known restoration suites). For a normal home recording, though, treating the room first and denoising lightly beats leaning on software. Our guide to recording a podcast at home covers capturing a clean take in the first place.

Carve problem frequencies with subtractive EQ

With a clean signal, reach for EQ — and think subtraction before addition. The most common spoken-word problem is a build-up of mud and boominess in the low mids. Sweep a narrow bell boost through roughly 150–400 Hz, find the frequency that sounds thick or “boxy”, then cut it a few dB with a wider Q. Proximity effect from close mic technique often piles energy here, so 2–4 dB of reduction around 200–300 Hz can be the single biggest clarity win.

Listen for a nasal or honky quality too, usually somewhere between 800 Hz and 1.5 kHz; a gentle 1–2 dB dip can open the voice up. Keep these cuts modest. Subtractive EQ on speech should be surgical and restrained — you’re removing distractions, not reshaping the voice. If you’re new to this, our primer on EQ and compression fundamentals explains the controls, and the Violet Recording free tools page has an EQ cheat sheet you can keep open while you work.

Compress for consistency, not character

Spoken delivery has a huge dynamic range — a whispered aside can be 15 dB quieter than an emphatic sentence. Compression is what glues that together so the listener never reaches for the volume knob. Start gentle: a ratio around 2:1 to 3:1, a medium attack of roughly 5–15 ms so consonants still punch through, and a medium release of around 60–150 ms that recovers between phrases. Aim for about 3–6 dB of gain reduction on the loudest lines.

For very dynamic voiceover, two light stages beat one heavy one. Use a first compressor to catch the biggest peaks, then a second doing another 2–3 dB — the result sounds more even and far less squashed than a single compressor working hard. This is the same principle behind our deeper walkthrough on mixing vocals. If your levels are wildly inconsistent before compression, some clip-gain riding by hand first will make the compressor’s job easier and more transparent. Getting your gain staging right at the input stage pays off here too.

Tame sibilance with a de-esser

Once you compress and add presence, “s”, “sh” and “t” sounds get louder and sharper — that’s sibilance, and it’s fatiguing over a long listen. A de-esser fixes it by ducking a narrow band only when those sounds occur. Set it to target roughly 5–9 kHz (higher voices sit toward the top of that range), and adjust the threshold so it clamps down just on the harsh sibilants, pulling maybe 3–6 dB on the worst ones. You should barely notice it working; if words start to lisp, you’ve gone too far. Your DAW’s stock de-esser handles this job well, so there’s rarely a need to buy one.

Add presence and a touch of saturation for intelligibility

Now build clarity back in with additive EQ. A gentle broad boost of 1–3 dB around 3–6 kHz brings out consonants and makes speech cut through on small speakers — this presence range is where intelligibility lives. A small lift above 10 kHz adds “air” and a professional, open quality, but use a shelf and keep it subtle so you don’t re-introduce sibilance. Always A/B against the unprocessed voice so the boosts stay honest.

A whisper of saturation can help too. Light harmonic drive thickens a thin voice and helps it feel present and forward without raising the fader. Stock saturation is fine for this; a tiny amount goes a long way on speech, so drive it until you can just hear it warm up, then back off. The aim throughout is a voice that sounds effortless and clear — never processed.

Level and hit a loudness target for the platform

The final stage is loudness. Spoken-word platforms publish integrated LUFS targets, and matching them keeps your programme from being turned down or sounding weak next to others. As approximate, widely used starting points: podcasts commonly aim for around -16 LUFS integrated (Apple Podcasts’ guideline), while music-streaming platforms normalise closer to -14 LUFS. Audiobook platforms like ACX ask for a tighter window — roughly -23 to -18 dB RMS with peaks below -3 dBTP and a low noise floor. Measure with a loudness meter rather than guessing, and check the exact spec for wherever you’re publishing. Our LUFS explainer breaks the numbers down, and the free tools page lists current platform targets.

To reach the target, use a limiter or a leveling plugin on the master, set a true-peak ceiling around -1 dBTP, and nudge the input until the meter lands in range — but don’t crush it, because over-limited speech sounds lifeless. If you’d rather hand off the loudness-and-consistency step, an online mastering service such as LANDR can normalise a spoken-word file to a chosen target automatically. For a full picture of that final stage, see what mastering actually does and browse the mixing and mastering hub for more technique guides.

Frequently asked questions

What’s different about mixing spoken word versus singing?

The priority flips from musicality to intelligibility. There’s no melody, harmony or beat to support the voice, so noise, boominess and sibilance are exposed and consistency matters more than tone-shaping. You’ll compress mainly for even level, keep EQ moves conservative, and pay close attention to plosives and background noise that a busy music mix would hide.

Do I need expensive plugins to mix spoken-word vocals?

No. Your DAW’s stock EQ, compressor, de-esser, saturation and limiter cover almost every step above, and a good recording matters far more than the brand of plugin. The one exception is heavy noise restoration on already-compromised recordings, where a dedicated tool clearly outperforms most built-in de-noisers. Otherwise, spend your effort on mic technique and room treatment rather than software.

How loud should a podcast or voiceover be?

Match the platform’s published loudness spec. A common practical target for podcasts is around -16 LUFS integrated with true peaks kept below about -1 dBTP, while audiobook distributors specify their own RMS and peak windows. Treat these as approximate starting points, measure with a loudness meter, and always confirm the current spec for your destination before you export.

Get the studio newsletter

New guides, gear deals and mixing tips — a couple of times a month. No spam, unsubscribe anytime.

More guides