Week 6 • Lesson 2 of 5 • 55 mins
AI Video: Clips and Avatars
Prompting generative clips, image-to-video for brands, and talking avatars with disclosure.
AI Video: Clips and Avatars
AI video generation has moved from blurry curiosities to usable footage in about two years. You can now turn a product photo into a moving shot, or record a message once and have an avatar deliver personalised versions to fifty clients.
It still has sharp limits — hands, crowds, complex physics, and consistency between shots. Working with those limits, not against them, is the whole skill.
Which model is best, how long a clip can be, and what it costs change almost monthly. Check the AI Tool Radar (
/tools) for current picks. The techniques below apply across tools.
1. Two different kinds of AI video
| Generative clips | Talking avatars | |
|---|---|---|
| What | New footage from a text prompt or an image | A person (real or synthetic) speaking a script with lip-sync |
| Tools | Runway, Kling, Google Veo, OpenAI Sora, Pika, Luma | HeyGen, Synthesia, D-ID |
| Good for | B-roll, product shots, ad visuals, mood scenes | Explainers, training, personalised outreach, translation |
| Weak at | Consistent characters, readable text, complex action | Emotion, gestures that match meaning, long takes |
2. Text-to-video vs image-to-video
Text-to-video invents everything. Great for mood and B-roll, unpredictable for anything that must match a brand.
Image-to-video starts from a picture you control — your product photo, your office, a designed frame — and adds motion. You keep the exact look, colours and composition.
Rule: if it must look like your product or brand, start from an image.
3. Prompting for video
A video prompt describes subject, action, camera, light, and style, in that order. Keep the action simple.
Subject: a ceramic coffee mug with steam rising, on a wooden table
Action: steam drifts slowly upward
Camera: slow push-in, shallow depth of field
Light: soft morning window light from the left
Style: photorealistic, warm tones, calm
Avoid: text, hands, extra objects, fast motion
Words that reliably help: slow pan, push-in, static shot, handheld, golden hour, shallow depth of field. Things that reliably go wrong: people eating, hands doing detailed work, crowds, anything with written text, fast complex motion.
4. Why you still need an editor
Each generation produces a short clip, and characters or objects drift between generations. A finished video is almost always:
- Several short generated shots
- Stitched in a normal editor (CapCut, DaVinci Resolve, Premiere, Descript)
- With your own voiceover or music, titles and captions added there — never generated inside the clip
Plan it like a shoot: write a shot list first, generate each shot, then edit.
5. Talking avatars: the workflow
- Script first. Write for the ear — short sentences, one idea each. Avatars can't rescue a bad script.
- Choose the avatar. Your own (record the training footage the platform asks for, well lit, looking at the lens) or a stock avatar.
- Voice. Your cloned voice (previous lesson) or a stock voice.
- Generate, then watch at full size. Check lip-sync on names and numbers.
- Edit in context. Cut in screen recordings or product shots every 10–20 seconds; a static talking head gets tiring fast.
Personalised outreach at scale: most avatar platforms can take a spreadsheet of first names and company names and produce one video per row. Keep personalisation to what you genuinely know — a fake-personal detail is worse than none.
6. Disclosure and consent
- Label avatar and synthetic video where a viewer would assume it's real. Several platforms (YouTube, Meta) require disclosure of realistic synthetic content, and India's IT Rules amendments in force since February 2026 require synthetically generated content on platforms to be clearly labelled.
- Never create an avatar of a real person without written consent.
- Never make a real person appear to say something they didn't. That's the definition of a harmful deepfake.
- Check whether you own commercial rights to generated footage on your plan before using it in ads.
⚠️ Common mistakes
- Complex action prompts. Keep motion simple: pans, zooms, drifting light.
- Expecting one generation to be a finished video. Plan shots, then edit.
- Text-to-video for branded products. Start from an image.
- Static avatar videos over a minute long. Cut away regularly.
- Undisclosed synthetic people. It's increasingly a platform rule, not just good manners.
What's next: from moving pictures to the format every professional still lives in — presentations.
Resources & Downloads
Hands-on Practicals
Use a free trial of HeyGen or D-ID. Upload a professional headshot and have it say: 'Hello, I am an AI-powered version of myself. I can speak 20 languages and I never sleep.' Share the video.
Create the same scene with 3 different prompt styles: 1) Basic description, 2) With cinematic terms (lighting, camera, style), 3) With negative prompts (what to avoid). Compare quality and specificity.
Create a short video explaining a concept using 5-10 second clips stitched together in CapCut. Use AI to generate the clips, then edit them into a coherent narrative. This is the future of rapid content creation.
Knowledge Check
Which tool is best for making a static photo of a person look like they are realistically speaking a script?
Why do most AI video tools require stitching clips together?
What is the benefit of using 'Image-to-Video' instead of 'Text-to-Video'?