The recording for this lesson has not been published yet.
In this lesson you produce natural-sounding English narration with controlled pacing and emphasis, then lock it to your captions in CapCut. The module principle still rules: script quality drives TTS quality more than any model setting. Whenever a line sounds wrong, ask first whether the text is the problem, and only then touch the sliders.
You will use the pipeline and naming convention from lesson 5.1. The English track you build here becomes the reference timing for the Polish version in lesson 5.3.
Choose the voice for the audience and format, not for how impressive it sounds in a 10-second demo. Audition every candidate with your script, not the sample text.
| Criterion | What to listen for | Red flag |
|---|---|---|
| Accent | Matches your audience (US, UK, neutral international) | Accent drifts between lines |
| Energy | Fits the format: calm for tutorials, brighter for promos | Hype voice on a serious topic |
| Clarity | Consonants clear on phone speakers | Breathy or mumbled endings |
| Consistency | Same character across 10+ renders | Each render sounds like a different person |
| Rights | Library voice license and plan allow your use | Cloned voice without documented consent |
If you use a cloned voice - even your client's own - keep written consent that covers commercial use and every language you will render. Do not imitate real people without permission. Check the ElevenLabs terms and your plan's license before delivery.
ElevenLabs exposes a few voice settings, typically sliders. Names, ranges and defaults differ between models and versions, so treat this table as a description of the kinds of controls you will meet, and always start from the defaults.
| Control | Lower values | Higher values | Starting point |
|---|---|---|---|
| Stability | More expressive and varied, but less predictable | More even and consistent, can become monotone | Default; raise for long tutorials, lower for storytelling |
| Similarity / clarity | Looser match to the original voice | Closer match; may reproduce artifacts from the source | Default to moderately high |
| Style / style exaggeration | Neutral delivery | Stronger character, less stable, slower renders on some models | Zero or low |
| Speed (if available) | Slower delivery | Faster delivery | Default; fix pacing in the script first |
| Speaker boost (if available) | Off | On: slightly closer to the voice, small latency cost | Try both, keep one |
With most models, the text itself is the prompt. You steer tone and rhythm with wording and punctuation:
<break time="0.5s" />; check the docs for your model, and do
not overuse it (too many can cause artifacts).FLAT
You can export in 4K and it only takes a minute and it keeps your captions.
TUNED
You can export in 4K. It takes about a minute.
And yes - your captions stay exactly where you put them.
| Sentence | Pace setting | Stability | Style strength | Notes |
|---|---|---|---|---|
| S01 hook: "Your edits are too slow." | Default | Medium-low | Low | Needs punch; split into two sentences |
| S02 explanation | Default | Medium-high | Zero | Calm, even, tutorial tone |
| S04 number reveal | Slightly slower | Medium | Low | Pause before the number with a full stop |
| S05 call to action | Default | Medium | Low | End on the action verb |
Proper nouns, brand names and acronyms are where English TTS most often fails. Fix them in this order: respell in the script, then use a pronunciation dictionary or alias feature if your plan and model support one, and only then consider re-recording with a different voice or model.
| Term | Intended pronunciation | Prompt hint (script respelling) | Final result |
|---|---|---|---|
| CapCut | CAP-cut | CapCut (usually fine) | OK first render |
| SQL | "sequel" | sequel | OK after respelling |
| GIF | Your house style, e.g. "gif" with hard g | ghif | OK, noted in style guide |
| API | A-P-I (letters) | A. P. I. or A-P-I | OK with dashes |
| Nguyen (person) | As the person says it - ask them | win | Approved by speaker |
| 1080p | ten eighty p | ten-eighty P | OK |
Respellings like "sequel" belong only in the TTS script. Your captions must show the real spelling (SQL). Keep a TTS version and a caption version of each line, or fix the captions after auto-generation.
Speeding up a VO clip in CapCut to make it fit changes pitch or adds artifacts and is easy to hear. If a line is too long for its shot, shorten the sentence and re-render, or extend the shot.