Week 6 • Lesson 1 of 5 • 50 mins

Voice AI: Cloning, Cleaning and Dubbing

Text-to-speech vs speech-to-speech, fixing recordings, dubbing — and the consent rules.

Voice AI: Cloning, Cleaning and Dubbing

AI voice is now good enough that most listeners can't tell a well-made clone from the real person. That makes it very useful — you can fix a flubbed sentence without re-recording, or publish a video in Hindi and Tamil in your own voice — and it makes it easy to misuse.

This lesson covers the three jobs voice AI does well, how to do each one properly, and the consent rules that aren't optional.

Tool features and plan limits in this area change monthly. The current recommended tools are on the AI Tool Radar (/tools). This lesson teaches the workflows, which stay stable.


1. The three jobs

Job What it does Typical tools Use it for
Text-to-speech (TTS) Types → voice ElevenLabs, built-in voices in Descript, Canva, CapCut Narration, explainers, IVR messages
Speech-to-speech (STS) Your recording → re-rendered voice, keeping your timing and emotion ElevenLabs Voice Changer Fixing a poor recording, a consistent brand voice
Dubbing Your video → other languages, in a voice like yours ElevenLabs Dubbing, HeyGen translation Reaching audiences in other languages

Plus one editing job:

  • Text-based audio editing — edit the transcript and the audio follows. In Descript, deleting "um" from the text removes it from the audio, and the feature once called Overdub (now part of its AI speech / Regenerate tools) lets you type a missing word and have it spoken in your cloned voice.

2. Text-to-speech vs speech-to-speech

This choice decides whether your audio sounds alive.

  • TTS invents all the delivery: pace, pauses, emphasis. Good TTS is pleasant but often flat over long passages.
  • STS keeps your human performance — where you paused, what you stressed, where you smiled — and only changes the voice quality.

Rule: if emotion matters (a story, a sales message, a personal update), record it yourself, even badly, and use speech-to-speech. Use TTS for neutral information.

3. Cloning your own voice properly

A clone is only as good as the sample.

  1. Quiet room, soft furnishings. Cupboards full of clothes are a free vocal booth.
  2. Consistent distance from the microphone — a fist's width for most USB and phone mics.
  3. Read naturally, in the tone you'll actually use. A clone trained on shouting sounds like shouting.
  4. Enough material. Instant clones work from about a minute; higher-quality clones want much more (often 30 minutes or more) of clean audio. Check your tool's current guidance.
  5. Test with a script you didn't train on, and listen on phone speakers, not just headphones.

Label AI-generated audio in descriptions where your audience would reasonably assume it's a live recording.

4. Consent is not optional

  • Only clone your own voice, or a voice whose owner has given explicit written permission for that specific use.
  • Reputable platforms require you to verify that a voice is yours before professional cloning — don't try to get around it.
  • Cloning someone else's voice without consent can be impersonation or fraud, and voice-cloned scam calls to family members are a real, growing crime. Agree a family "safe word" for urgent money requests.
  • For staff voices in company content, get consent in writing, including what happens to the clone when they leave.

5. Workflow: fix a recording without re-recording

  1. Record your video or podcast as normal.
  2. Import into a text-based editor (e.g. Descript).
  3. Remove filler words in bulk, then listen back — over-removing makes speech sound breathless.
  4. Fix a wrong word: correct the transcript and regenerate just that phrase in your voice.
  5. Studio sound / noise removal on the whole track, at a moderate setting.
  6. Export and listen once, start to finish, before publishing.

6. Workflow: dub a video into another language

  1. Start with clean audio. Background music confuses dubbing — keep a version without it.
  2. Get the transcript right first. Names, product terms and numbers wrong in the source will be wrong in every language.
  3. Run the dub, then have a native speaker check at least the first minute. Machine translation handles meaning; it often misses register (formal vs casual) and local terms.
  4. Export subtitles too. Many viewers watch muted.
  5. Publish each language as its own upload or audio track, with a note that it's AI-dubbed.

7. Music on demand, carefully

Tools like Suno and Udio generate full songs from a prompt:

Upbeat instrumental, 90 seconds, acoustic guitar and light
percussion, warm and optimistic, suitable as a podcast intro,
no vocals, clean ending.

Before commercial use, read the licence for your plan — rights to generated music differ between free and paid tiers, and the legal position on AI music is still being tested in court. For ads, a stock library with a clear licence is often the safer choice.


⚠️ Common mistakes

  • Monotone TTS for emotional content. Use speech-to-speech instead.
  • Cloning from a noisy sample. Garbage in, robot out.
  • Cloning anyone else's voice without explicit written permission.
  • Skipping the native-speaker check on dubbed content.
  • Using AI music commercially without reading the licence.

What's next: your audio sounds professional. Now the visual side — generating video clips and talking avatars.

Hands-on Practicals

The Voice Clone

Record yourself reading a 1-minute story. Use ElevenLabs to clone your voice and have it read a completely different story. How close was the match?

Speech-to-Speech Comparison

Record yourself speaking naturally (good and bad takes). Use Speech-to-Speech to clean it up. Compare: 1) Naturalness, 2) Emotion, 3) Time saved vs. re-recording. This demonstrates the power of preserving human performance.

Dubbing Workflow

Record a 2-minute video explaining a concept. Use ElevenLabs dubbing to translate it into 3 languages while keeping your voice. Create subtitle files for each version. This is how you scale content globally.

Knowledge Check

What is 'Speech-to-Speech' technology?

What is the primary ethical consideration when using voice cloning?

What does Descript's Overdub (now part of its AI speech / Regenerate tools) let you do?