Course resource
AI Music Prompt Pack
Generating usable music and audio, and the licensing rules that decide whether you can actually publish it.
The prompt structure
[GENRE], [MOOD], [INSTRUMENTATION], [TEMPO], [VOCALS OR INSTRUMENTAL], [USE CASE]
Weak: upbeat background music
Working:
Warm acoustic folk, optimistic but understated, fingerpicked guitar with
light brushed drums and upright bass, around 95 BPM, instrumental,
for a two-minute product explainer video — must sit under a voiceover
without competing with it
The use case is the part people omit and it matters most: music for a voiceover needs different dynamics from music for a title sequence.
Ready-made starting points
Under a voiceover
Sparse instrumental bed, [GENRE], no melody in the vocal range,
steady dynamics with no sudden swells, [TEMPO] BPM, loopable,
mixed to sit at background level
Intro or title sequence
[GENRE], confident opening hit within the first second, builds over
8 seconds to a clean resolve, [INSTRUMENTATION], no fade-in
Podcast theme
[GENRE] theme, memorable 4-bar motif, [MOOD], [TEMPO] BPM,
arranged so it can be cut at 5, 15 and 30 seconds without sounding truncated
Ambient work music
Ambient [GENRE], no vocals, no percussion, no melodic hooks,
minimal variation, 20+ minutes, designed not to draw attention
With vocals
[GENRE], [MALE / FEMALE / DUET] vocals, [VOCAL STYLE], [MOOD],
lyrics about [THEME], chorus that repeats [HOOK LINE], [TEMPO] BPM
The controls that matter
| Element | What it changes | Ranges worth knowing |
|---|---|---|
| Tempo | Energy, more than genre does | 60–80 calm · 90–110 conversational · 120–140 energetic |
| Instrumentation | Character | Naming 3–4 instruments beats naming a genre |
| Mood | Emotional register | Use two words that qualify each other: "optimistic but restrained" |
| Structure | Usability | Say where it builds, where it resolves, whether it loops |
| Mix note | Whether it works in context | "sits under dialogue", "no low end", "wide stereo" |
Getting a usable take
Generation is cheap and inconsistent. Expect to generate ten and use one.
- Generate 8–10 on the same prompt.
- Listen only to the first 10 seconds of each. Discard ruthlessly.
- Take the two best and regenerate variations on those.
- Check the full length of the finalists — many fall apart after 45 seconds.
- Check it loops, if you need it to.
Licensing — read this before publishing
This is where people get hurt.
- Free tiers usually do not grant commercial rights. Generating on a free plan and putting it on a monetised video is a licence breach, regardless of the tool.
- Rights differ by tier and change over time. Check the current terms of your specific plan, not a blog post.
- Some platforms retain the right to reuse your generations. If exclusivity matters, check.
- Content ID is a real risk. AI-generated music has been matched against Content ID systems, sometimes wrongly. Keep your generation records.
- Do not prompt with an artist's name for anything commercial. "In the style of [living artist]" is a legal and ethical problem.
Keep a record for everything you publish:
| Track | Tool | Plan/tier | Date | Prompt | Licence permits commercial? |
|---|---|---|---|---|---|
If a claim ever lands, this table is what resolves it.
Voice and narration
The same structural approach applies to text-to-speech. What decides quality:
- Punctuation is direction. Commas, full stops and paragraph breaks control pacing more than any setting.
- Spell out anything ambiguous. Numbers, acronyms, and names — write "twenty twenty-six" if the reading matters.
- Break long text into paragraphs. Quality degrades across very long single generations.
- Speech-to-speech beats text-to-speech when you want real emotion — record yourself badly, let the model re-render it. Your timing and emphasis survive.
Voice cloning has its own rules — see the Voice Cloning Ethics guidelines. The short version: never clone a voice without written consent.
Before you publish
- My plan's licence permits this use
- I have not named a living artist in the prompt
- I have listened to the whole thing, not just the opening
- It loops cleanly, if it needs to
- It does not fight the voiceover
- Generation record saved
- Disclosed, if the audience would care