Week 8 • Lesson 4 of 5 • 40 mins
Detecting AI Content — and Its Limits
Why detector scores aren't verdicts, weak and strong signals, and a fair verification protocol.
Detecting AI Content — and Its Limits
Teachers, editors, hiring managers and clients all want to know: was this written by AI? Detection tools promise an answer with a percentage.
The honest position: no current tool can reliably tell you whether a specific piece of text was written with AI help. They're useful as one signal among several, and harmful as a verdict.
1. What the detection tools actually do
Tools like GPTZero, Originality.ai, Turnitin's AI indicator and Copyleaks estimate how statistically "predictable" text is, compared with patterns typical of model output.
What independent studies consistently find:
- Overall accuracy is moderate — studies report figures roughly in the 60–85% range depending on the tool and text type. Not definitive.
- Mixed human + AI writing is close to undetectable. A human draft polished with AI, or an AI draft heavily edited by a human, defeats them.
- False positives are real and uneven. Writing by non-native English speakers is flagged as AI far more often than native speakers' writing — some studies found the majority of genuine essays by non-native writers flagged. Formulaic genres (technical, legal, scientific) are also over-flagged.
- Light paraphrasing of AI text often drops detection sharply.
- OpenAI withdrew its own AI text classifier in 2023, citing low accuracy.
So a "92% AI" score is not a 92% probability that someone cheated. Accusing someone on that basis is unfair and, in education and employment, can cause real harm.
2. The "AI-isms" — weak signals
Some patterns are common in unedited model output:
- Stock transitions: "Furthermore", "Moreover", "Additionally", "In conclusion"
- Inflated vocabulary: "delve", "multifaceted", "intricate", "tapestry", "pivotal", "landscape"
- Structure fetish: perfectly balanced headings and bullets; every paragraph the same length
- Hedge-and-summarise: "It's important to note that…", a tidy recap at the end of everything
- Generic examples with no specific names, places, dates or numbers
- Occasional leftovers: "As an AI…", "Certainly! Here's…"
These are weak signals. Many humans write this way (especially in corporate writing), and newer models and simple instructions remove them. Never treat a word like "delve" as proof.
3. Human markers — better signals
Authentic human writing tends to contain things a model can't supply without the person:
- Specific lived detail: "the client in Nashik who paid in cash every third month"
- Consistent personal voice across a person's work — their usual phrases, rhythm, punctuation habits, humour
- Honest imperfection: tangents, unusual phrasing, a strong opinion not balanced away
- Knowledge of context that wasn't in any prompt: the internal meeting, the class discussion
4. The verification protocol
Instead of a detector verdict, when authenticity genuinely matters:
- Compare voice with the person's known previous work.
- Check specifics: are claims, quotes and citations real and verifiable? Fabricated references are a far stronger signal than style.
- Talk to them. Ask them to explain their reasoning, a choice they made, or to expand on a point. Understanding (or its absence) shows quickly.
- Look at process evidence: drafts, version history, notes.
- Use a detector score only as a prompt to look closer, never as evidence on its own, and never as the sole basis for a decision.
5. Better than detecting: clear rules
The real fix is upstream:
- Say what's allowed. "AI may be used for outlining and grammar; the analysis must be yours; disclose how you used it."
- Design for it. Oral defences, in-class writing, process portfolios, personal and local topics.
- Ask for disclosure as normal practice, without penalty for honest, permitted use.
6. Synthetic images, audio and video
For media, the signals are different:
- Provenance data: check for Content Credentials (C2PA) labels and platform "AI-generated" tags — many generators and cameras now attach them, though they can be stripped.
- Visual clues: inconsistent hands, text, jewellery, reflections, lighting direction; backgrounds that melt.
- Audio clues: unnatural breathing, flat emotion on emotional words, odd pronunciation of names.
- Context checks: reverse image search; does a reputable source carry it? Would this person plausibly say this?
For any urgent request for money or credentials by voice or video — even from a familiar face or voice — verify through a separate channel you already trust.
⚠️ Common mistakes
- False certainty from a detector percentage.
- Accusing someone — especially a non-native English writer — based on a tool score.
- Treating "AI-isms" as proof.
- Paranoia that treats all AI-assisted work as dishonest.
- Missing the point: the goal is clear rules and genuine understanding, not a perfect detector.
What's next: pulling it all together — a governance framework for yourself or your team that people will actually follow.
Resources & Downloads
Hands-on Practicals
Take 5 pieces of content (your own, AI-generated, AI-assisted). Try to identify which is which using the human touch markers. Check with detection tools. Compare your accuracy vs. the tool accuracy. This builds your intuition.
Collect 3 writing samples from the same person spanning different times/ contexts. Ask AI to generate a 'fake' sample in that person's style. Can you tell the difference? Can detection tools? This is the cutting edge of the detection arms race.
Use 3 different AI detection tools on the same set of 10 samples (mixtures of human, AI, and AI-assisted). Calculate each tool's accuracy. Document which tool performed best for your use case.
Knowledge Check
What is the current accuracy range of AI detection tools?
What is a reliable method to detect AI-assisted writing?
Which of these is NOT a reliable approach to AI detection?