Video coming soon

The recording for this lesson has not been published yet.

What we cover

  • Noise reduction basics
  • Loudness consistency
  • EQ and clarity for speech
  • Monitoring with headphones

Lesson notes

1. What this lesson is for

In this lesson you turn raw spoken audio into clean, even, speech-first sound that stays clear on a phone speaker, on earbuds and on studio headphones. Keep the module principle in front of you: intelligible speech is non-negotiable. Music, effects and style come later (lesson 4.2) - none of them can rescue a voice the viewer cannot understand.

By the end of the module you will deliver one finalized mix where dialogue or voiceover stays clear across phone and headphone playback. This lesson builds the first half of that mix: cleanup, leveling and EQ.

2. The mixing flow

Always work in the same order. Each step depends on the one before it: denoise changes the level, leveling changes how EQ sounds, and ducking only works if the voice is already consistent.

Cleanup Leveling EQ Music / SFX layering Ducking Loudness check Solid boxes: this lesson. Dashed boxes: lesson 4.2. The loudness check is repeated at the very end of the mix.

3. Noise reduction basics

Noise is anything that is not the voice: room hum, fans, traffic, hiss from a cheap mic. CapCut offers a noise reduction toggle (sometimes with a strength setting, sometimes just on/off) in the audio panel of a selected clip. The exact name and options differ between desktop, mobile and web versions, so look for Reduce noise or similar under the clip's audio settings.

  • Fix it at the source first. Closer mic, quieter room, soft furnishings. No plugin beats 20 cm less distance to the mouth.
  • Denoise before anything else. Leveling and compression raise the noise floor together with quiet speech, so remove noise while it is still low.
  • Less is more. Use the lightest setting that makes the noise unobtrusive, not absent. A little room tone sounds natural; zero noise with a metallic voice sounds broken.
  • Listen to the pauses and the consonants. Artifacts show up first in gaps (a "watery" swirl) and on S, T, F sounds (a lisp or phasing).

4. Cleanup chain template

Use this as your default chain for a single voice clip. Parameter ranges are starting points; CapCut may expose only some of these controls (or none, as an on/off switch) in your version. If a control is missing, do the step in a dedicated audio tool and re-import.

Step Tool Parameter range Warning signs
1. Trim silence and clicksSplit / trim, short fadesFades of 3-10 ms at edit pointsClicks or pops at cuts
2. Noise reductionReduce noiseLowest strength that worksMetallic, watery or lisping voice
3. Low cut (high-pass)EQ70-100 Hz, gentle slopeThin voice if set above ~120 Hz
4. Clip gain / normalizeVolume, loudness normalizationSpeech peaks around -6 to -3 dBFSSome clips obviously louder than others
5. Presence / clarityEQ+1 to +3 dB around 2-5 kHzHarsh, tiring or sibilant voice
6. De-ess (if available)De-esser or narrow EQ cut-2 to -6 dB around 5-8 kHzLisp, dull S sounds
7. Final ceilingLimiter / export normalizationTrue peak -1 dBTPDistortion on loud words, clipping

5. Loudness consistency

Two numbers matter. Peak is the loudest single moment; it tells you about clipping. Loudness (measured in LUFS, Loudness Units relative to Full Scale) is the perceived average; it tells you whether your video sounds as loud as the one before it in the feed. Platforms normalize playback toward a target, so being far louder than the target buys you nothing and costs you dynamics.

0 dB -1 -6 -14 -30 peak ceiling -1 dBTP integrated loudness ~ -14 LUFS leveled voice: peaks stay between -6 and -3, never touch the ceiling unleveled: spikes clip, quiet words vanish

Reference targets per platform (approximate; platforms change these without notice):

Platform Integrated loudness True peak Notes
YouTubeabout -14 LUFS-1 dBTPLouder uploads are turned down; quieter ones are generally not turned up
TikTok / Instagram Reels / Shorts-14 to -12 LUFS-1 dBTPNo official public spec; mostly phone playback, so speech clarity beats loudness
Spotify / podcasts-14 LUFS (Spotify), -16 LUFS (common podcast)-1 to -2 dBTPRelevant if you reuse the audio as a podcast
Broadcast (EBU R128)-23 LUFS-1 dBTPOnly if you deliver to TV

And the per-track reference for the whole module (the mix you finish in lesson 4.2):

Track type Target level Peak ceiling Notes
Dialogue / voiceoverLoudest element, peaks -6 to -3 dBFS-1 dBTPAlways on top; everything else is measured relative to it
Music under speech15-25 dB below voice-6 dBFSShould feel present, never compete
Music without speech6-10 dB below voice level-3 dBFSIntros, transitions, breaks
SFX accentsShort peaks near voice level-3 dBFSBrief; avoid masking words
Full mixabout -14 LUFS integrated-1 dBTPMeasure after export

CapCut does not always show a LUFS meter. It may offer a loudness normalization option on clips or at export; if it does not, measure the exported file with a free tool such as ffmpeg:

# measure integrated loudness (I), true peak (TP) and range (LRA)
ffmpeg -i final_mix.mp4 -af loudnorm=print_format=summary -f null -

# typical output to read
Input Integrated:    -17.8 LUFS    -> about 4 dB too quiet for -14
Input True Peak:      -0.2 dBTP    -> too hot, aim for -1.0
Input LRA:             6.1 LU      -> fine for speech (roughly 4-10)

Log every clip while leveling so you can see what changed and why:

LOUDNESS LOG - episode_03
CLIP            PEAK     AVERAGE    GAIN APPLIED   RESULT
intro_vo        -1.5     -20 LUFS   -2 dB          peaks now -3.5, ok
interview_a     -9.0     -26 LUFS   +5 dB          matches intro
interview_b     -4.0     -22 LUFS   +1 dB          ok, slight hiss audible
street_take     -0.1     -16 LUFS   -4 dB + denoise  was clipping, now clean
outro_vo        -6.0     -21 LUFS   +1 dB          ok
FULL MIX        -1.0 dBTP  -14.2 LUFS  -           pass

6. EQ and clarity for speech

EQ shapes the tone of the voice. CapCut exposes EQ as presets or a simple band editor depending on the version; the frequency guide below applies to any EQ.

Range What lives there Typical move Too much sounds like
Below 80 HzRumble, handling noise, trafficCut (high-pass)-
100-250 HzBody and warmthLeave, or small cut if boomyCut: thin; boost: muddy
250-500 HzBoxiness, "room"-2 to -4 dB narrow cutCut too much: hollow
2-5 kHzPresence, intelligibility of consonants+1 to +3 dB wide boostHarsh, tiring
5-8 kHzSibilance (S, SH)Cut or de-ess if sharpCut too much: lisp
Above 10 kHzAirSmall shelf boost if dullHiss becomes obvious

7. Monitoring with headphones and phone

Laptop speakers hide both bass problems and hiss. Checking the mix only on laptop speakers is one of this module's listed pitfalls. Use this routine every time:

  1. Headphones (closed-back if possible) at a moderate, fixed volume - hunt for noise, clicks, denoise artifacts and sibilance.
  2. Phone speaker - export or preview on the phone. Is every word understandable without subtitles?
  3. Earbuds - check that nothing is harsh or tiring at high volume.
  4. Log the differences, then fix and re-check.
MONITORING LOG - episode_03
TIME      DEVICE       ISSUE                          FIX
00:12     headphones   hiss in pause                  denoise one step up
00:34     phone        "fifteen" sounds like "fifty"  +2 dB at 3 kHz on interview_a
00:51     earbuds      sharp S on "system"            de-ess / narrow cut at 6.5 kHz
01:05     phone        outro quieter than intro       +1.5 dB clip gain

8. QA checklist

  • Noise floor reduced without metallic or watery artifacts.
  • Levels normalized - no clip is noticeably louder or quieter than its neighbours.
  • No clipping: true peak at or below -1 dBTP on the export.
  • Consonants remain crisp; no harsh sibilance.
  • Voice keeps its body - EQ cuts did not make it thin.
  • Checked on headphones and a phone speaker, with differences logged.

9. Exercise

  1. Record or pick a 60-90 second voice clip with some room noise and uneven volume. Duplicate the project so you keep the original.
  2. Apply the cleanup chain from section 4 in order. Write down each setting you used.
  3. Fill in a loudness log for every clip, then export and measure the full mix (target about -14 LUFS, -1 dBTP).
  4. Run the monitoring routine on headphones and a phone speaker and log at least three differences.
  5. Retrospective: where did denoise introduce a metallic sound, and where did an EQ cut reduce the body of the voice? What would you change at recording time?

Further reading

After this lesson you will

  • Reduce noise and even out voice volume
  • Prepare clear spoken audio for social platforms