The recording for this lesson has not been published yet.
In this lesson you turn raw spoken audio into clean, even, speech-first sound that stays clear on a phone speaker, on earbuds and on studio headphones. Keep the module principle in front of you: intelligible speech is non-negotiable. Music, effects and style come later (lesson 4.2) - none of them can rescue a voice the viewer cannot understand.
By the end of the module you will deliver one finalized mix where dialogue or voiceover stays clear across phone and headphone playback. This lesson builds the first half of that mix: cleanup, leveling and EQ.
Always work in the same order. Each step depends on the one before it: denoise changes the level, leveling changes how EQ sounds, and ducking only works if the voice is already consistent.
Noise is anything that is not the voice: room hum, fans, traffic, hiss from a cheap mic. CapCut offers a noise reduction toggle (sometimes with a strength setting, sometimes just on/off) in the audio panel of a selected clip. The exact name and options differ between desktop, mobile and web versions, so look for Reduce noise or similar under the clip's audio settings.
If the voice starts to sound robotic, underwater or "like a phone call", the reduction is too strong. Back it off, or accept a bit of noise. Aggressive denoise artifacts are one of this module's listed pitfalls - viewers forgive hiss far more easily than a warbling voice.
Use this as your default chain for a single voice clip. Parameter ranges are starting points; CapCut may expose only some of these controls (or none, as an on/off switch) in your version. If a control is missing, do the step in a dedicated audio tool and re-import.
| Step | Tool | Parameter range | Warning signs |
|---|---|---|---|
| 1. Trim silence and clicks | Split / trim, short fades | Fades of 3-10 ms at edit points | Clicks or pops at cuts |
| 2. Noise reduction | Reduce noise | Lowest strength that works | Metallic, watery or lisping voice |
| 3. Low cut (high-pass) | EQ | 70-100 Hz, gentle slope | Thin voice if set above ~120 Hz |
| 4. Clip gain / normalize | Volume, loudness normalization | Speech peaks around -6 to -3 dBFS | Some clips obviously louder than others |
| 5. Presence / clarity | EQ | +1 to +3 dB around 2-5 kHz | Harsh, tiring or sibilant voice |
| 6. De-ess (if available) | De-esser or narrow EQ cut | -2 to -6 dB around 5-8 kHz | Lisp, dull S sounds |
| 7. Final ceiling | Limiter / export normalization | True peak -1 dBTP | Distortion on loud words, clipping |
Two numbers matter. Peak is the loudest single moment; it tells you about clipping. Loudness (measured in LUFS, Loudness Units relative to Full Scale) is the perceived average; it tells you whether your video sounds as loud as the one before it in the feed. Platforms normalize playback toward a target, so being far louder than the target buys you nothing and costs you dynamics.
Reference targets per platform (approximate; platforms change these without notice):
| Platform | Integrated loudness | True peak | Notes |
|---|---|---|---|
| YouTube | about -14 LUFS | -1 dBTP | Louder uploads are turned down; quieter ones are generally not turned up |
| TikTok / Instagram Reels / Shorts | -14 to -12 LUFS | -1 dBTP | No official public spec; mostly phone playback, so speech clarity beats loudness |
| Spotify / podcasts | -14 LUFS (Spotify), -16 LUFS (common podcast) | -1 to -2 dBTP | Relevant if you reuse the audio as a podcast |
| Broadcast (EBU R128) | -23 LUFS | -1 dBTP | Only if you deliver to TV |
And the per-track reference for the whole module (the mix you finish in lesson 4.2):
| Track type | Target level | Peak ceiling | Notes |
|---|---|---|---|
| Dialogue / voiceover | Loudest element, peaks -6 to -3 dBFS | -1 dBTP | Always on top; everything else is measured relative to it |
| Music under speech | 15-25 dB below voice | -6 dBFS | Should feel present, never compete |
| Music without speech | 6-10 dB below voice level | -3 dBFS | Intros, transitions, breaks |
| SFX accents | Short peaks near voice level | -3 dBFS | Brief; avoid masking words |
| Full mix | about -14 LUFS integrated | -1 dBTP | Measure after export |
CapCut does not always show a LUFS meter. It may offer a loudness normalization option on
clips or at export; if it does not, measure the exported file with a free tool such as
ffmpeg:
# measure integrated loudness (I), true peak (TP) and range (LRA)
ffmpeg -i final_mix.mp4 -af loudnorm=print_format=summary -f null -
# typical output to read
Input Integrated: -17.8 LUFS -> about 4 dB too quiet for -14
Input True Peak: -0.2 dBTP -> too hot, aim for -1.0
Input LRA: 6.1 LU -> fine for speech (roughly 4-10)
Log every clip while leveling so you can see what changed and why:
LOUDNESS LOG - episode_03
CLIP PEAK AVERAGE GAIN APPLIED RESULT
intro_vo -1.5 -20 LUFS -2 dB peaks now -3.5, ok
interview_a -9.0 -26 LUFS +5 dB matches intro
interview_b -4.0 -22 LUFS +1 dB ok, slight hiss audible
street_take -0.1 -16 LUFS -4 dB + denoise was clipping, now clean
outro_vo -6.0 -21 LUFS +1 dB ok
FULL MIX -1.0 dBTP -14.2 LUFS - pass
EQ shapes the tone of the voice. CapCut exposes EQ as presets or a simple band editor depending on the version; the frequency guide below applies to any EQ.
| Range | What lives there | Typical move | Too much sounds like |
|---|---|---|---|
| Below 80 Hz | Rumble, handling noise, traffic | Cut (high-pass) | - |
| 100-250 Hz | Body and warmth | Leave, or small cut if boomy | Cut: thin; boost: muddy |
| 250-500 Hz | Boxiness, "room" | -2 to -4 dB narrow cut | Cut too much: hollow |
| 2-5 kHz | Presence, intelligibility of consonants | +1 to +3 dB wide boost | Harsh, tiring |
| 5-8 kHz | Sibilance (S, SH) | Cut or de-ess if sharp | Cut too much: lisp |
| Above 10 kHz | Air | Small shelf boost if dull | Hiss becomes obvious |
Removing mud (250-500 Hz) often makes a voice clearer than boosting presence. If an EQ cut makes the voice lose its body, you cut too low or too wide - move the band up or narrow it.
Laptop speakers hide both bass problems and hiss. Checking the mix only on laptop speakers is one of this module's listed pitfalls. Use this routine every time:
MONITORING LOG - episode_03
TIME DEVICE ISSUE FIX
00:12 headphones hiss in pause denoise one step up
00:34 phone "fifteen" sounds like "fifty" +2 dB at 3 kHz on interview_a
00:51 earbuds sharp S on "system" de-ess / narrow cut at 6.5 kHz
01:05 phone outro quieter than intro +1.5 dB clip gain
Pushing the voice up until the meter hits 0 dB causes distortion and is then turned down by the platform anyway. Aim for consistent, clean speech at the target, not maximum volume.