Course resource

AI Video Prompt Book

Getting usable clips out of a video generator, and knowing when not to bother.

The prompt structure

[SHOT TYPE] of [SUBJECT] [ACTION], [SETTING], [LIGHTING],
[CAMERA MOVEMENT], [STYLE], [DURATION AND PACING]

Weak: a person walking in a city

Working:

Medium tracking shot of a woman in a grey raincoat walking briskly
along a wet pavement, neon-lit street at night, reflections in puddles,
camera follows from her left at walking pace, cinematic, shallow depth
of field, single continuous 5-second shot, no cuts

Camera movement is the highest-leverage term

Naming the movement is what makes output look filmed rather than generated.

Term Effect
Static / locked off Calm, observational
Slow push in Building attention
Pull back / reveal Context, surprise
Tracking shot Energy, following
Pan left/right Surveying a space
Handheld Immediacy, documentary
Drone / aerial Scale
Orbit Showcasing an object

Say "single continuous shot, no cuts" on almost everything. Generators will invent cuts, and an unintended cut mid-clip is the most common reason a take is unusable.

What still fails

Be realistic about what you will get. Current generators are strong on short, simple, physically plausible shots and weak on everything else.

The working method: generate many short clips, discard most, edit the survivors together. Do not try to generate the finished video.

Prompts by purpose

B-roll

[SHOT TYPE] of [SUBJECT], [SETTING], [LIGHTING], slow [MOVEMENT],
no people's faces, 4 seconds, single continuous shot, cinematic,
neutral colour grade for editing

Product

Slow orbit around [PRODUCT] on [SURFACE], soft studio lighting from
above and left, seamless [COLOUR] background, shallow depth of field,
macro detail on [FEATURE], 5 seconds, no text, no hands

Establishing shot

Wide aerial of [LOCATION] at [TIME OF DAY], [WEATHER], slow forward
drone movement, natural colour, 6 seconds, single continuous shot

Talking-head avatar

Script: [YOUR SCRIPT]
Framing: medium close-up, eyes to camera
Background: [PLAIN / OFFICE / BLURRED]
Delivery: [CONVERSATIONAL / AUTHORITATIVE], moderate pace

For avatars, the script matters far more than the visual prompt. Write for the ear: short sentences, no subordinate clauses, punctuation where you want breath.

The workflow that produces usable video

  1. Storyboard first, in text. List every shot you need before generating anything.
  2. Generate 5–10 takes per shot. Assume a 1-in-8 hit rate.
  3. Judge on the first second. If the first second is wrong, the rest will be.
  4. Watch finalists at full speed twice. Artifacts hide in motion — hands, background people, cloth, reflections.
  5. Cut in a real editor. The generator makes clips; the edit makes the video.
  6. Add text, logos and captions afterwards. Never generate them.
  7. Grade the whole thing together so clips from different generations match.

Cost control

Video generation is the most expensive thing in this course by an order of magnitude.

Before you publish

When not to use it

Back to dashboard