Week 2 • Lesson 2 of 6 • 60 mins
Multimodal AI: Vision, Voice, and Video
Using AI to analyze images, clone voices, and generate videos without coding.
The Multimodal Revolution: Beyond Text
For two years, AI meant "chat with text." That era is over. Modern AI can see, hear, speak, and create across every media format. This isn't future tech - it's available right now, and it's reshaping how creative work gets done.
1. Vision AI - Teaching Machines to See
What's actually happening: Vision AI models can analyze images at a pixel level and at a semantic level. They understand both "there's a red object at coordinates X,Y" and "this is a sunset photo with warm emotional tone."
What this reliably does today:
- Object detection and classification
- Text extraction (OCR) from images
- Scene understanding and description
- Style and aesthetic analysis
- Comparison across multiple images
- Basic reasoning about visual information
What it CAN'T do (yet):
- Perfectly count small objects (still makes mistakes)
- Understand very fine detail in complex diagrams
- "See" things that aren't visually present (can't read minds from photos)
Practical Use Cases:
Business Use Case 1: Competitive Analysis
[Upload competitor's advertisement]
"Analyze this ad. Tell me:
1. Color psychology - what emotions are they targeting ?
2. Visual hierarchy - where does the eye go first ?
3. Messaging approach - what's the core promise?
4. Target demographic - who is this for?
5. Design trends they're using
6. What makes this effective or ineffective ?
Be specific.Reference actual elements of the ad."
You get professional-grade creative analysis in seconds.
Business Use Case 2: Brand Consistency Audit
[Upload 5 - 10 images from your brand's social media]
"Analyze these images for brand consistency:
- Color palette: What colors appear ? Are they consistent ?
- Style : What's the visual style? Photography, illustration, 3D?
- Composition: Any patterns in how elements are arranged ?
- Tone : Professional, playful, minimalist, bold ?
- Identify outliers - which images don't match the others?
Create a brand guideline document based on the patterns you see."
Your brand guidelines, reverse-engineered from existing content.
Daily Life Use Case: Recipe from Ingredients
[Upload photo of your refrigerator contents]
"I have these ingredients. Give me 3 dinner recipes:
- One quick(under 30 minutes)
- One healthy(under 500 calories)
- One impressive(could serve to guests)
For each recipe:
- List what I'll use from the photo
- What else I'll need (if anything)
- Step - by - step instructions
- Approximate time and difficulty"
Meal planning without the mental load.
Creative Use Case: Design Feedback
[Upload your design mockup]
"As a senior UX designer, critique this interface:
- Visual hierarchy: Is the most important element most prominent ?
- White space: Too cramped or too sparse ?
- Color accessibility: Any contrast issues ?
- Alignment and consistency
- Three specific improvements I should make
Be direct.I need actionable feedback, not compliments."
Professional design review without hiring a consultant.
Advanced Vision Techniques:
Multi-image comparison:
[Upload your before / after photos]
"Compare these two images. What changed? Be specific about:
- Colors and lighting
- Composition and framing
- Mood and emotional impact
- Technical quality
- Which is more effective for [specific purpose]"
Visual data extraction:
[Upload screenshot of a table or chart]
"Extract this data into a formatted table. Then:
- Identify the 3 key trends
- Note any anomalies or surprises
- Suggest what additional data would be useful"
Style transfer analysis:
[Upload two images]
"Image 1 is our current brand aesthetic. Image 2 is a style we're considering. Analyze:
- What visual elements would change ?
- What stays the same ?
- How difficult would this transition be ?
- What's the risk of alienating existing audience?"
Platform-specific tips:
ChatGPT Vision (GPT-4V/4o):
- Best for general image analysis
- Great OCR for extracting text from images
- Can generate images (DALL-E) in the same conversation
- Good at reading charts, screenshots, diagrams
Claude Vision:
- Better at detailed analytical work
- Stronger reasoning about what it sees
- More conservative (less likely to hallucinate details)
- Can handle longer, more complex visual analysis
Gemini Vision:
- Can analyze video (upload video files)
- Integrated with Google Search (can fact-check visual claims)
- Good at multi-image comparison
- Strong at document and PDF analysis
Common Vision AI Mistakes:
Mistake #1: Assuming perfect accuracy Vision AI is good but not infallible. It can:
- Miscount objects (especially small ones)
- Misidentify people or things
- Miss subtle details
- Make assumptions about context
Fix: Verify critical details, especially for business decisions.
Mistake #2: Low-quality images Blurry, dark, or low-resolution images produce poor analysis.
Fix: Use clear, well-lit, high-resolution images.
Mistake #3: Vague requests "Analyze this image" gets you a generic description.
Fix: Be specific: "Analyze the color scheme and explain the psychological impact of each color choice."
Mistake #4: Expecting mind-reading AI can only analyze what's visible. It can't tell you the photographer's intent or hidden meanings.
Fix: Ask about what's visually present, not what's implied.
2. Voice AI - The Sound of Silicon
Voice AI has hit an inflection point. We're past robotic text-to-speech. Current voice AI is eerily human.
The landscape:
Text-to-Speech (TTS) - AI reads text aloud:
ElevenLabs (Gold Standard)
- Most natural-sounding voices
- Voice cloning from 30 seconds of audio
- Emotional control (you can direct tone)
- Multiple languages
- Professional quality
Use cases:
- Podcast narration (create content without recording)
- Video voiceovers
- Audiobook creation
- Accessibility (turn your blog into audio versions)
Pricing:
- Free tier: 10k characters/month (about 10 minutes of audio)
- Paid: $5-$99/month depending on volume
OpenAI TTS (Built into ChatGPT)
- Simpler but solid
- Multiple voice options
- Lower cost
- Good for basic needs
Play.ht / Murf.ai (Professional alternatives)
- More voice options
- Better editing tools
- Team collaboration features
Voice Cloning - Your digital voice twin:
How it works:
- Record 30-60 seconds of your voice reading provided text
- AI analyzes: pitch, tone, cadence, accent, speech patterns
- Model created that can "speak" as you
Ethical use:
- Clone your own voice for content creation
- With explicit permission, clone someone else (like a client testimonial)
Unethical use:
- Cloning without permission
- Impersonation for scams
- Deceptive use
Reality check: Most platforms require you to confirm "I have rights to clone this voice" before proceeding. They're trying to prevent misuse.
Practical applications:
Use Case 1: Consistent brand voice across content
Clone your voice → Use it for:
- YouTube video narration(consistent across all videos)
- Podcast episodes(even when you're sick)
- Course content and tutorials
- Social media voice notes
Benefit: Your audience hears "you" even when you didn't personally record.
Use Case 2: Multilingual content
Some voice AI tools can:
- Clone your voice
- Have it speak other languages
- Maintain your voice characteristics
This is wild - your voice, speaking fluent Spanish or French, despite you not speaking those languages.
Limitation: Quality varies by language.English is best, major languages good, rare languages sketchy.
Use Case 3: Rapid content production
Write a 2000 - word blog post → TTS → Audio version in 5 minutes
Traditional way:
- Set up mic
- Find quiet space
- Record(multiple takes for mistakes)
- Edit audio
- Export
Total: 1 - 2 hours
AI way:
- Paste text
- Select voice
- Generate
Total: 2 minutes
Music Generation - AI Composers:
Suno AI (Current Leader)
- Creates complete songs from text prompts
- Lyrics + melody + instruments + production
- Multiple genres
- Surprisingly professional quality
Example prompt:
"Create an indie folk song about leaving corporate life to travel. Male vocals, acoustic guitar, gentle percussion. Melancholic but hopeful tone. 3 minutes long."
Output: Complete, radio-ready song.
Udio (Strong Alternative)
- Similar to Suno
- Some say better at specific genres
- Different audio characteristics
What you can create:
- Background music for videos
- Podcast intros/outros
- Hold music for business phone systems
- Custom songs for events (birthday, wedding)
- Jingles for brands
Limitations:
- Copyright is murky (check latest ToS for commercial use)
- Can't consistently generate specific melodies (it's creative, not precise)
- Voice quality varies (sometimes uncanny valley)
- Limited fine control (you can't say "make the bridge more upbeat")
Best for:
- Background music where original composition isn't critical
- Prototyping (create temp music, then hire musician for final version)
- Personal projects
- Low-budget content
NotebookLM Audio Overview - The Secret Weapon:
This Google tool is slept on and it's incredible.
What it does: Uploads PDFs, docs, notes → Generates a 10-15 minute podcast where two AI hosts discuss your content
Why it's amazing:
- Two voices (male and female) having natural conversation
- They explain concepts, give examples, connect ideas
- Sounds like an NPR podcast, not robots
- Great for learning - you can listen during commute
Use cases:
Academic:
Upload: Research papers on quantum computing
Output: Podcast where two "hosts" break down the papers, explain technical concepts, debate interpretations
Listen to it while exercising instead of reading dense PDFs.
Business:
Upload: Your company's product docs, strategy memos, market research
Output: Podcast briefing for new team members or executives
Onboard people faster - they listen to a 15 - minute show instead of reading 50 pages.
Personal Learning:
Upload: Articles you saved on a topic you're learning
Output: Synthesized discussion that connects the ideas
Passive learning while doing dishes or commuting.
Voice AI Mistakes:
Mistake #1: Using default voices The generic robot voices make content feel cheap. Invest 30 minutes to find or clone a good voice.
Mistake #2: No script editing AI voices sound best with well-written scripts. If your text is full of run-on sentences or unclear phasing, the audio will sound weird.
Fix: Edit for the ear, not the eye. Read your script aloud before generating audio.
Mistake #3: Ignoring emotional direction Many TTS tools let you specify emotion (excited, serious, sad). Using monotone when content is exciting kills engagement.
Mistake #4: Copyright ignorance Especially with AI music - know the licensing terms. Some platforms grant commercial rights, others don't.
3. Video AI - The Visual Revolution
Video generation is where things get really wild. We're in the early days, but it's advancing fast.
Current capabilities:
Runway Gen-3 (Industry Leader)
- Text-to-video: Describe a scene, get a 5-10 second clip
- Image-to-video: Upload an image, it animates it
- Video-to-video: Transform existing video's style
- Professional cinematic quality
Pika Labs (Strong Alternative)
- Similar capabilities to Runway
- Different aesthetic (some prefer it for certain styles)
- Good at camera movements and effects
What you can create:
- B-roll for videos (landscape shots, abstract visuals)
- Product showcase animations
- Social media content
- Concept visualization
- Artistic/experimental content
Practical examples:
Marketing video B-roll:
Prompt: "Aerial shot of a modern city at golden hour, camera slowly pushing forward, cinematic lighting, 4K quality"
Result: 10 seconds of professional - looking city footage
Use for: Background in product videos, website headers, social posts
Product visualization:
Upload: Still image of your product
Prompt: "Slowly rotate 360 degrees, studio lighting, clean white background"
Result: Animated product showcase
Use for: E - commerce listings, social media, presentations
Concept pitch:
Prompt: "A futuristic office where AI and humans collaborate, warm lighting, optimistic tone, people working alongside holographic displays"
Result: Visual concept for pitch deck
Use for: Investor presentations, internal vision documents
Talking Head Videos - HeyGen and D-ID:
These tools create "AI avatars" - digital people who speak your script.
HeyGen (Most Popular)
- Upload a photo (or use stock avatar)
- Paste your script
- AI animates the photo to lip-sync perfectly
- Multiple languages with automatic translation
Use cases:
Training videos:
Instead of filming yourself explaining 50 different topics:
- Write scripts for each topic
- Generate avatar videos
- Consistent look, professional quality, no re - filming
Multilingual content:
Record yourself explaining something in English
- HeyGen can create versions where "you" speak Spanish, French, Hindi, etc.
- Your face, perfect lip - sync, translated speech
Expand global reach without hiring translators and voice actors.
Quick announcements:
Company update video needed today:
- Write the script
- Generate avatar video
- Publish
Total time: 15 minutes instead of half - day video shoot
Realistic expectations (what sucks right now):
Video AI limitations:
- Consistency: Hard to generate the same character/object across multiple clips
- Length: Most tools max out at 10-15 seconds per clip
- Physics: Sometimes objects move strangely or defy physics
- Text: Terrible at generating text in scenes (it's gibberish)
- Complex motion: Struggles with detailed human actions
- Cost: Professional tools are expensive ($30-100/month)
What this means:
- Great for supplemental content, B-roll, concepts
- Not ready to replace human creators for most use cases
- But improving rapidly - watch this space
Legal and ethical considerations:
Deepfakes and consent: Using someone's likeness without permission = unethical and potentially illegal
Commercial rights:
- Check if AI-generated content can be used commercially
- Some platforms retain rights, others grant full ownership
- Especially important for client work or products you sell
Disclosure: Many platforms now add watermarks to AI-generated video. Some contexts require disclosing content is AI-made.
4. Putting It All Together - Multimodal Workflows
The real power is combining these tools:
Complete Content Pipeline:
1. Text(ChatGPT): Write script for explainer video
2. Voice(ElevenLabs): Generate narration from script
3. Visuals(Runway / Midjourney): Create supporting visuals
4. Video(Editing tool): Combine everything
5. Result: Professional explainer video, zero filming
Research to Podcast Pipeline:
1. Vision(Claude): Analyze research papers and visual data
2. Text(ChatGPT): Synthesize findings into a report
3. Voice(NotebookLM): Generate podcast discussing findings
4. Result: Audio briefing of complex research
Brand Refresh Pipeline:
1. Vision AI: Analyze current brand visuals
2. Text AI: Generate brand guidelines based on analysis
3. Image AI(Midjourney): Create new brand visual examples
4. Present: Complete brand refresh proposal
What's Next:
You now understand each modality. But these are tools - powerful tools that need structure and systems.
Next lesson: Building your personal AI assistant with Custom GPTs and Claude Projects. Creating specialized tools that know your context, your voice, and your workflow.
Resources & Downloads
Hands-on Practicals
Create a 3-slide presentation on any topic without using traditional design tools. Process: (1) Use ChatGPT to write the content for each slide, (2) Use Midjourney or DALL-E to generate images for each slide, (3) Use ElevenLabs to create narration for each slide, (4) Combine in Google Slides or PowerPoint. Time yourself - this should take under 30 minutes. Compare to how long it would normally take to create, design, and record a presentation. Track: What was easy? What was frustrating? Where did you need to iterate?
Upload 5-7 images from your brand's social media (or a brand you admire if you don't have one) to Claude or ChatGPT Vision. Ask for a comprehensive brand analysis covering: color psychology, visual patterns, photography style, composition patterns, tone and mood, target demographic signals. Then ask it to identify the 1-2 images that don't quite fit. Finally, have it create a one-page brand guideline based on the patterns. Use this as the foundation for actual brand guidelines.
Record a 60-second voice sample of yourself reading neutral text (read a Wikipedia article aloud). Upload to ElevenLabs and clone your voice. Then test it: Have your cloned voice read (1) a professional business announcement, (2) a casual social media script, (3) technical instructions. Listen critically: Does it capture your natural tone? Where does it sound robotic? What emotions transfer well? What breaks the illusion? Most people are shocked how good it is, but notice specific words or phrases that sound off.
Pick a topic you need to learn (work-related or personal interest). Find 3-5 substantial articles or PDFs on that topic. Upload all of them to NotebookLM. Generate the Audio Overview podcast. Listen to it once, then: (1) Write down the key concepts the 'hosts' discussed, (2) Identify what questions you still have, (3) Ask NotebookLM those follow-up questions, (4) Listen to the audio again after asking questions. Track: How much did you retain vs. reading the PDFs yourself? Was the audio format more or less effective for this type of content?
Find a complex infographic, chart, or table online (or screenshot one from a report). Upload to ChatGPT Vision and ask it to: (1) Extract all data into a formatted markdown table, (2) Identify the 3 key trends, (3) Note any surprising data points, (4) Suggest 2 different ways to visualize this data. Verify accuracy by checking the extracted data against the source. Track error rate - this helps you calibrate when to trust vision AI for data work vs. manual entry.
Knowledge Check
You photograph a printed table of quarterly figures and ask AI to extract it into a spreadsheet. What should you do before using the numbers?
What is the primary advantage of Suno/Udio over traditional music creation?
What makes NotebookLM's Audio Overview feature particularly valuable for learning?
What is a critical limitation to be aware of when using Vision AI for business decisions?