The recording for this lesson has not been published yet.
What we cover
Choosing non-distracting background music
Sound effects for emphasis
Manual and automatic ducking
Fade-in/out and timing transitions
Lesson notes
1. What this lesson is for
In this lesson you layer music and sound effects under a clean voice track (from lesson
4.1) so they add emotion and emphasis without ever costing a word. The module principle
still rules: intelligible speech is non-negotiable. If music and voice
fight, the voice wins every time.
This completes the module deliverable: one finalized mix where dialogue or voiceover stays
clear across phone and headphone playback.
2. The layer map: who has priority
Think of your mix as three layers stacked by importance. Every decision about level,
timing and EQ follows this order.
3. Choosing non-distracting background music
Instrumental under speech. Lyrics compete directly with narration for
the same attention. Save vocal tracks for sections without speech.
Sparse midrange. Speech lives roughly between 300 Hz and 4 kHz. Pads,
soft piano, light percussion and bass leave room; busy guitars, synth leads and brass
fill that space.
Steady energy. Pick tracks without huge drops or builds in the middle
of a spoken section, or cut around them.
Match the tempo to the edit. Energetic content suits 110-130 BPM;
calm explainers suit 70-95 BPM. Cuts on the beat feel intentional.
Check the licence. CapCut's built-in library is licensed for use in
CapCut, but terms for commercial or off-platform use vary by track and region. Read the
track's licence notes in your version before publishing.
Carve a hole for the voice
If a track is almost right but still masks words, apply a gentle EQ cut of 2-4 dB around
1-4 kHz on the music only. The voice pops through without turning the music down further.
4. Sound effects for emphasis
SFX guide attention: a whoosh sells a transition, a pop marks a text reveal, a click
confirms a UI action. Use them sparingly - one per idea, not one per cut.
SFX
Visual action
Timing
Mistake to avoid
Whoosh / swish
Transition, fast pan, slide-in
Peak lands on the cut frame
Whoosh on every cut
Pop / bubble
Text or sticker appearing
Starts on the first visible frame
Covering the first word of a line
Click / tap
Screen recording, button press
Exactly on the press
Being 3+ frames late
Riser
Build-up before a reveal
Ends on the reveal
Running under narration
Hit / impact
Key number, punch-in
On the emphasis frame
Too loud, startling the viewer
Room tone / ambience
Scene setting
Continuous, very low
Being louder than the music
Place SFX on their own track so you can adjust them without touching the music. Zoom in and
align the waveform transient (the sharp spike) with the visual frame, not the start of the
audio clip.
5. Manual and automatic ducking
Ducking means lowering the music while someone speaks and bringing it back
up in the gaps. The shape of that level change over time is the ducking envelope.
Manual ducking with keyframes. Select the music clip, add a volume
keyframe just before the voice starts (full level), another at the voice start (ducked
level), and mirror them at the end. Four keyframes per spoken section give you full
control. This works in every CapCut version that has audio keyframes.
Automatic ducking. Some CapCut versions offer an automatic ducking or
"auto duck" option on music tracks that lowers them wherever voice is detected. Where it
exists, its name, location and adjustable parameters differ between desktop and mobile,
so check your build. Use it for a fast first pass, then fix problem spots manually.
Split-and-lower (quick fallback). Split the music at the start and end
of the speech, lower the middle piece and add short fades. Crude, but it works when
keyframes are fiddly on mobile.
Avoid pumping
Pumping is the music audibly jumping up in every short pause between sentences. It
happens when release is too fast or auto ducking reacts to each breath. Bridge gaps
shorter than about 1 s - keep the music ducked through them - and lengthen the release.
6. Ducking plan template
Plan levels before you draw keyframes. Values are relative to your clip volume in CapCut
(0 dB = unchanged), assuming the voice is already leveled as in lesson 4.1.
Segment
Voice level
Music level
Duck amount
Fade times
Intro, no speech
-
-8 dB
0 dB
Fade in 1-2 s
Energetic narration
0 dB
-22 dB
about 14 dB
Attack 0.2 s, release 0.6 s
Calm explainer
0 dB
-26 dB
about 18 dB
Attack 0.4 s, release 1.2 s
Transition / B-roll break
-
-10 dB
0 dB
Swell up over 0.5-1 s
Key statement
0 dB
silence or -30 dB
full
Fade out 0.5 s before the line
Outro
-
-8 dB
0 dB
Fade out 2-4 s
DUCKING PLAN - episode_03 (music: calm_piano_82bpm.mp3)
TIME SEGMENT MUSIC KEYFRAMES
00:00-00:04 intro 0.0 s: -40 dB -> 1.5 s: -8 dB
00:04-00:31 narration (calm) 3.6 s: -8 dB -> 4.0 s: -26 dB
31.0 s: -26 dB -> 32.2 s: -10 dB
00:32-00:36 B-roll break hold -10 dB
00:36-00:58 narration 35.6 s: -10 dB -> 36.0 s: -26 dB
00:48 key statement 47.5 s: -26 dB -> 48.0 s: -40 dB (silence)
00:58-01:04 outro 58.0 s: -8 dB -> 64.0 s: -40 dB
7. Fade-in/out and timing transitions
Every music clip gets a fade. Even 0.3 s at the start and end removes
clicks. CapCut exposes fade-in and fade-out durations on selected audio clips.
Cut music on the beat or phrase. If you shorten a track, join the two
pieces on a downbeat and cross-fade over a few frames so the edit is invisible.
End with the music, not against it. Trim the track so its natural
ending lines up with your last frame, or fade out over 2-4 s.
Use silence. Dropping the music completely just before a key line is
one of the strongest emphasis tools you have.
8. QA checklist
Voice always dominant - every word understandable on a phone speaker.
Ducking transitions smooth, with no pumping between sentences.
SFX cues align with visual actions (within 1-2 frames).
No clipping on the export: true peak at or below -1 dBTP, full mix about -14 LUFS.
Smooth fades between sections; no clicks at music edits.
Checked on headphones and a phone, not only laptop speakers.
Music too loud under narration
This is the most common mixing mistake. You know the script, so you understand the words
even when the music masks them - your viewer does not. When in doubt, turn the music
down another 3 dB.
9. Exercise
Take the cleaned voice track from lesson 4.1 and duplicate the project.
Mix two segments: one energetic (faster music, stronger SFX) and one calm (sparse
music, few or no SFX), each with a different music intensity profile.
Write a ducking plan like the one above, then build it with volume keyframes. If your
version has automatic ducking, compare it with your manual result.
Add three SFX tied to visual actions and align them frame-accurately.
Export, measure loudness, and check on headphones and a phone speaker.
Review: where does music still compete with speech, where would silence improve
impact, and what could you simplify or remove?