Video coming soon

The recording for this lesson has not been published yet.

What we cover

  • Selecting an English voice profile
  • Prompting for tone and pacing
  • Handling names, acronyms and technical words
  • Syncing English voiceover with captions

Lesson notes

1. What this lesson is for

In this lesson you produce natural-sounding English narration with controlled pacing and emphasis, then lock it to your captions in CapCut. The module principle still rules: script quality drives TTS quality more than any model setting. Whenever a line sounds wrong, ask first whether the text is the problem, and only then touch the sliders.

You will use the pipeline and naming convention from lesson 5.1. The English track you build here becomes the reference timing for the Polish version in lesson 5.3.

2. Selecting an English voice profile

Choose the voice for the audience and format, not for how impressive it sounds in a 10-second demo. Audition every candidate with your script, not the sample text.

Criterion What to listen for Red flag
AccentMatches your audience (US, UK, neutral international)Accent drifts between lines
EnergyFits the format: calm for tutorials, brighter for promosHype voice on a serious topic
ClarityConsonants clear on phone speakersBreathy or mumbled endings
ConsistencySame character across 10+ rendersEach render sounds like a different person
RightsLibrary voice license and plan allow your useCloned voice without documented consent

3. Voice settings: what the controls do

ElevenLabs exposes a few voice settings, typically sliders. Names, ranges and defaults differ between models and versions, so treat this table as a description of the kinds of controls you will meet, and always start from the defaults.

Control Lower values Higher values Starting point
StabilityMore expressive and varied, but less predictableMore even and consistent, can become monotoneDefault; raise for long tutorials, lower for storytelling
Similarity / clarityLooser match to the original voiceCloser match; may reproduce artifacts from the sourceDefault to moderately high
Style / style exaggerationNeutral deliveryStronger character, less stable, slower renders on some modelsZero or low
Speed (if available)Slower deliveryFaster deliveryDefault; fix pacing in the script first
Speaker boost (if available)OffOn: slightly closer to the voice, small latency costTry both, keep one
low stability high stability target zone for narration expressive, emotional random emphasis, odd pitch jumps even, predictable flat, robotic over long reads Change one control at a time and render the same line to compare.

4. Prompting for tone and pacing

With most models, the text itself is the prompt. You steer tone and rhythm with wording and punctuation:

  • Pauses - commas, full stops, and paragraph breaks. An ellipsis often gives a hesitant pause; a dash gives a short break. Some models also accept a break tag such as <break time="0.5s" />; check the docs for your model, and do not overuse it (too many can cause artifacts).
  • Emphasis - put the key word at the end of a short sentence, where stress naturally lands. Capitals sometimes add stress but are unreliable.
  • Tone - context words shape delivery. A question mark rises; an exclamation adds energy. Newer expressive models may support tone tags in brackets - only use them if your model's docs list them.
  • Pace - short sentences read faster and punchier; long ones slow the listener down.
FLAT
You can export in 4K and it only takes a minute and it keeps your captions.

TUNED
You can export in 4K. It takes about a minute.
And yes - your captions stay exactly where you put them.
Sentence Pace setting Stability Style strength Notes
S01 hook: "Your edits are too slow."DefaultMedium-lowLowNeeds punch; split into two sentences
S02 explanationDefaultMedium-highZeroCalm, even, tutorial tone
S04 number revealSlightly slowerMediumLowPause before the number with a full stop
S05 call to actionDefaultMediumLowEnd on the action verb

5. Handling names, acronyms and technical words

Proper nouns, brand names and acronyms are where English TTS most often fails. Fix them in this order: respell in the script, then use a pronunciation dictionary or alias feature if your plan and model support one, and only then consider re-recording with a different voice or model.

Term Intended pronunciation Prompt hint (script respelling) Final result
CapCutCAP-cutCapCut (usually fine)OK first render
SQL"sequel"sequelOK after respelling
GIFYour house style, e.g. "gif" with hard gghifOK, noted in style guide
APIA-P-I (letters)A. P. I. or A-P-IOK with dashes
Nguyen (person)As the person says it - ask themwinApproved by speaker
1080pten eighty pten-eighty POK

6. QC routine: listen twice

  1. Listen at 1x with the script in front of you. Mark wrong words, odd stress and missing pauses.
  2. Listen at 1.25x without the script. Robotic phrasing and unnatural pauses become much more obvious when sped up.
  3. For each flagged line decide: script rewrite (the wording fights the rhythm) or parameter tweak (the wording is fine, the delivery is off). Rewrites fix most problems.
  4. Re-render only the flagged lines, save as the next version, and log the change.

7. Syncing English voiceover with captions

  • Place the approved VO clips first; captions follow the audio, never the reverse.
  • Use CapCut's auto-captions on the VO track, then correct spelling against the caption script (real spellings, not respellings).
  • Keep caption blocks to one short line or two at most, and break them at natural pauses, not mid-phrase.
  • Scrub the timeline at key words: the caption should appear with, or a few frames before, the spoken word.
  • When a VO clip is replaced, recheck all captions after it.

8. Checklist

  • Proper nouns and acronyms rendered correctly and logged.
  • Emphasis lands on the words that carry the message.
  • Pauses are intentional, not random.
  • Caption spelling matches the real terms, and timing matches the audio.
  • Voice, model and settings recorded for future re-renders.

9. Exercise

  1. Audition three English voices with the first two lines of your script and pick one using the criteria table.
  2. Render the full script, then run the two-speed QC routine and fill in the pronunciation tracker.
  3. For one line, try three settings variations (change one control each time) and note which you keep and why.
  4. Rewrite at least two lines for rhythm instead of changing settings, and compare.
  5. Sync the VO and captions in CapCut and export a test video.
  6. Retrospective: which lines needed a script rewrite versus a parameter tweak?

Further reading