Video coming soon

The recording for this lesson has not been published yet.

What we cover

  • Choosing non-distracting background music
  • Sound effects for emphasis
  • Manual and automatic ducking
  • Fade-in/out and timing transitions

Lesson notes

1. What this lesson is for

In this lesson you layer music and sound effects under a clean voice track (from lesson 4.1) so they add emotion and emphasis without ever costing a word. The module principle still rules: intelligible speech is non-negotiable. If music and voice fight, the voice wins every time.

This completes the module deliverable: one finalized mix where dialogue or voiceover stays clear across phone and headphone playback.

2. The layer map: who has priority

Think of your mix as three layers stacked by importance. Every decision about level, timing and EQ follows this order.

1. Voiceover / dialogue - always dominant 2. SFX accents - short, tied to visual actions 3. Music bed - foundation, mood, pacing priority When two layers collide, lower the lower layer - never raise the upper one past its ceiling.

3. Choosing non-distracting background music

  • Instrumental under speech. Lyrics compete directly with narration for the same attention. Save vocal tracks for sections without speech.
  • Sparse midrange. Speech lives roughly between 300 Hz and 4 kHz. Pads, soft piano, light percussion and bass leave room; busy guitars, synth leads and brass fill that space.
  • Steady energy. Pick tracks without huge drops or builds in the middle of a spoken section, or cut around them.
  • Match the tempo to the edit. Energetic content suits 110-130 BPM; calm explainers suit 70-95 BPM. Cuts on the beat feel intentional.
  • Check the licence. CapCut's built-in library is licensed for use in CapCut, but terms for commercial or off-platform use vary by track and region. Read the track's licence notes in your version before publishing.

4. Sound effects for emphasis

SFX guide attention: a whoosh sells a transition, a pop marks a text reveal, a click confirms a UI action. Use them sparingly - one per idea, not one per cut.

SFX Visual action Timing Mistake to avoid
Whoosh / swishTransition, fast pan, slide-inPeak lands on the cut frameWhoosh on every cut
Pop / bubbleText or sticker appearingStarts on the first visible frameCovering the first word of a line
Click / tapScreen recording, button pressExactly on the pressBeing 3+ frames late
RiserBuild-up before a revealEnds on the revealRunning under narration
Hit / impactKey number, punch-inOn the emphasis frameToo loud, startling the viewer
Room tone / ambienceScene settingContinuous, very lowBeing louder than the music

Place SFX on their own track so you can adjust them without touching the music. Zoom in and align the waveform transient (the sharp spike) with the visual frame, not the start of the audio clip.

5. Manual and automatic ducking

Ducking means lowering the music while someone speaks and bringing it back up in the gaps. The shape of that level change over time is the ducking envelope.

Voice narration Music level -10 -28 attack 0.2-0.5 s release 0.5-1.5 s duck amount ~15-20 dB music bed held low under speech keyframe
  • Manual ducking with keyframes. Select the music clip, add a volume keyframe just before the voice starts (full level), another at the voice start (ducked level), and mirror them at the end. Four keyframes per spoken section give you full control. This works in every CapCut version that has audio keyframes.
  • Automatic ducking. Some CapCut versions offer an automatic ducking or "auto duck" option on music tracks that lowers them wherever voice is detected. Where it exists, its name, location and adjustable parameters differ between desktop and mobile, so check your build. Use it for a fast first pass, then fix problem spots manually.
  • Split-and-lower (quick fallback). Split the music at the start and end of the speech, lower the middle piece and add short fades. Crude, but it works when keyframes are fiddly on mobile.

6. Ducking plan template

Plan levels before you draw keyframes. Values are relative to your clip volume in CapCut (0 dB = unchanged), assuming the voice is already leveled as in lesson 4.1.

Segment Voice level Music level Duck amount Fade times
Intro, no speech--8 dB0 dBFade in 1-2 s
Energetic narration0 dB-22 dBabout 14 dBAttack 0.2 s, release 0.6 s
Calm explainer0 dB-26 dBabout 18 dBAttack 0.4 s, release 1.2 s
Transition / B-roll break--10 dB0 dBSwell up over 0.5-1 s
Key statement0 dBsilence or -30 dBfullFade out 0.5 s before the line
Outro--8 dB0 dBFade out 2-4 s
DUCKING PLAN - episode_03   (music: calm_piano_82bpm.mp3)
TIME          SEGMENT              MUSIC KEYFRAMES
00:00-00:04   intro                0.0 s: -40 dB -> 1.5 s: -8 dB
00:04-00:31   narration (calm)     3.6 s: -8 dB  -> 4.0 s: -26 dB
                                   31.0 s: -26 dB -> 32.2 s: -10 dB
00:32-00:36   B-roll break         hold -10 dB
00:36-00:58   narration            35.6 s: -10 dB -> 36.0 s: -26 dB
00:48         key statement        47.5 s: -26 dB -> 48.0 s: -40 dB (silence)
00:58-01:04   outro                58.0 s: -8 dB  -> 64.0 s: -40 dB

7. Fade-in/out and timing transitions

  • Every music clip gets a fade. Even 0.3 s at the start and end removes clicks. CapCut exposes fade-in and fade-out durations on selected audio clips.
  • Cut music on the beat or phrase. If you shorten a track, join the two pieces on a downbeat and cross-fade over a few frames so the edit is invisible.
  • End with the music, not against it. Trim the track so its natural ending lines up with your last frame, or fade out over 2-4 s.
  • Use silence. Dropping the music completely just before a key line is one of the strongest emphasis tools you have.
fade in ducked under voice swell silence: key line fade out

8. QA checklist

  • Voice always dominant - every word understandable on a phone speaker.
  • Ducking transitions smooth, with no pumping between sentences.
  • SFX cues align with visual actions (within 1-2 frames).
  • No clipping on the export: true peak at or below -1 dBTP, full mix about -14 LUFS.
  • Smooth fades between sections; no clicks at music edits.
  • Checked on headphones and a phone, not only laptop speakers.

9. Exercise

  1. Take the cleaned voice track from lesson 4.1 and duplicate the project.
  2. Mix two segments: one energetic (faster music, stronger SFX) and one calm (sparse music, few or no SFX), each with a different music intensity profile.
  3. Write a ducking plan like the one above, then build it with volume keyframes. If your version has automatic ducking, compare it with your manual result.
  4. Add three SFX tied to visual actions and align them frame-accurately.
  5. Export, measure loudness, and check on headphones and a phone speaker.
  6. Review: where does music still compete with speech, where would silence improve impact, and what could you simplify or remove?

Further reading

After this lesson you will

  • Layer music and effects under narration
  • Use ducking so speech remains clear