Course Insights

    The Vizzini Dilemma of AI Detection

    Why schools and students are both trapped in circular reasoning, and how to escape.

    Vizzini spiraling through AI detection logic while Turnitin sits calmly across the table

    The Princess Bride is one of my favourite films. Any excuse to use it to make a point, I will take. So here we are.

    There's a scene where the criminal Vizzini faces Wesley in a battle of wits. There are two wine glasses, and one is poisoned. Vizzini has to choose. He launches into this elaborate chain of reasoning: "You must have suspected I would have known the powder's origin, so I can clearly not choose the wine in front of you. But you must have known I was not a great fool, so you would have counted on it, so I can clearly not choose the wine in front of me." Round and round. "I know that you know that I know."

    This is exactly how schools are approaching AI detection right now. Students try to avoid tells. Detectors get updated. Students update their evasion strategies. Detectors update again. Everyone's trapped in the same unwinnable guessing game.

    The thing about Vizzini's reasoning is it's actually sound given the premises of the game. The problem is that he has committed one of the classic blunders, only slightly less famous than "Never start a land war in Asia": he accepted the game at all.

    The Evidence (The Arms Race Is Unwinnable)

    The entire detection process is structurally flawed. The evidence for this comes from researchers, the AI creators themselves, Nordic education ministries, and elite US institutions, all reaching the same conclusion independently.

    Temple University tested Turnitin against 120 sample texts. For hybrid texts (human writing mixed with AI, which is how students actually work), accuracy was only 23-57%. The sentences Turnitin flagged had "no relationship at all" to the sections that were actually AI-generated. The "evidence" shown to teachers is essentially random.

    OpenAI discontinued their own AI Classifier in July 2023. It correctly identified AI text only 26% of the time. It falsely flagged human text 9% of the time. If the creators of ChatGPT can't build a reliable detector, no one can.

    A Stanford study tested seven popular AI detectors on TOEFL essays versus US 8th-grade writing. 61% of non-native English essays were falsely flagged as AI-generated. The detectors were near-perfect for native speakers. The mechanism is straightforward: non-native speakers write with limited vocabulary, standard structures, and clarity, exactly what AI also produces. This matters in IB schools full of EAL students who would be disproportionately harmed.

    False Positive Rates by Demographics

    AI detectors disproportionately harm non-native English speakers

    Native English Writers
    0%
    falsely flagged as AI
    Non-Native (TOEFL) Writers
    61%
    falsely flagged as AI

    Systematic Discrimination

    EAL students face a 61% false positive rate while native English speakers face 0%. This isn't a bug—it's discrimination baked into the technology.

    Source: Stanford University Study (2023) on GPT detectors

    This bias matters. Second language speakers are much more likely to write in the patterns that AI also produces: limited vocabulary, standard structures, clarity. These aren't tells for cheating. They're characteristics of developing fluency. The 0% false positive rate for native speakers isn't proof the detectors work. It's proof they're systematically biased against EAL students.

    Norway's Directorate for Education advised schools against using AI detectors, explicitly stating they are "not reliable." Their response was to remove the preparation phase of exams and block internet access during testing, a retreat from digital surveillance back to analogue verification. Director Morten Rosenkvist's guidance focused on ensuring "it is you, and not artificial intelligence, who has written the answer."

    MIT Sloan issued guidance titled "AI Detectors Don't Work." They recommend faculty focus on policy clarity and relationship-building rather than surveillance.

    The pattern is clear. Detection is broken.

    Data simulated based on research regarding adversarial attacks on NLP models

    The "False Positive" Trap

    When institutions rely on unreliable tools, they create a cycle of distrust. Experts warn that the damage to the student-teacher relationship outweighs the benefit of catching a few cheaters.

    ✍️
    Student Submits Work
    ⚠️
    Detector Flags False Positive
    (e.g., due to ESL writing style)
    👮
    Accusation Made
    Burden of proof on student
    💔
    Trust Broken
    Education environment harmed

    Why We Can't Just Stop Assigning Essays

    The lazy response is "just don't assign take-home work." That dodges the real problem.

    AI is an amplifier, not a replacement. It multiplies what you bring to it. Bring weak thinking, you get polished weak thinking. Bring nothing, you get coherent emptiness. Bring genuine understanding, you get genuinely useful output.

    The essay used to be reasonable evidence that cognitive work had happened because producing it was hard. That link is now broken. Detection tries to restore it by inferring process from product, but it can't.

    What broke wasn't "students are cheating more." The assumption that product implies process was always a shortcut.

    Michael Sandel, in what I feel is his most compelling book, The Tyranny of Merit, writes about how meritocratic systems develop a kind of moral dependence on their own metrics. Grades, credentials, outputs — they become load-bearing. University admission and high school exams are some of the most entrenched meritocratic systems, and their moral dependence on those metrics runs deep. Grades stop being evidence of learning and start being the thing itself.

    When AI undermines the essay as a reliable signal, the institutional reflex is to authenticate the signal rather than question whether we were measuring the right thing at all. Detection is the system trying to protect its own scaffolding. This isn't new; contract cheating — paying someone to do the work for you — exploited the same gap long before AI existed. AI just democratised it.

    Why Do They Fail? The Mechanics

    Detectors don't "know" what is written. They look for mathematical patterns.

    1. Perplexity: A measure of randomness. Low randomness (predictable text) is flagged as AI. High randomness (creative/chaotic text) is human.
    2. Burstiness: Variation in sentence structure. AI tends to be monotonous. Humans vary sentence length.

    "The problem is that academic writing is taught to be predictable, structured, and clear. Exactly the traits detectors flag as 'AI'."

    The Wesley Move (Process Over Product)

    Wesley wins the battle of wits not by participating in Vizzini's mental acrobatics but by zooming out. He knows that whichever glass you choose, poison bad. His strategy from the start is to be immune to poison.

    The educational equivalent is simple for teachers; it is to stop trying to infer process from product.

    What this looks like in practice:

    • Oral defenses and vivas (the IB already uses these)
    • Version history in documents (working in Google Docs where edit history is preserved)
    • Staged submissions (proposal, outline, draft, final)
    • Knowing your students' voices and patterns

    The IB's own guidance points this direction. Emphasizing teacher judgment, oral follow-ups, and process documentation rather than relying on detection tools.

    The question shifts from "Did AI write this?" to "Can this student defend this?"

    Synthesized estimation based on trends from major universities (Vanderbilt, Michigan, MIT, etc.)

    Wesley's immunity strategy

    None of this means schools are acting in bad faith. The pressure to verify authenticity is real, and a tool that promises to solve it with a percentage score is enormously tempting. I work in a school. I sit in meetings where these tools get discussed seriously by thoughtful people. The frustration is genuine — something changed, and doing nothing feels irresponsible. I get that. The problem is that the reassurance detection offers is hollow. The numbers don't hold up, and the students most likely to be harmed are the ones already working hardest.

    For students

    Be Wesley, Not Vizzini

    Don't get trapped in the loop of "how do I evade detection?" That's playing the wrong game. The students gaming detection and the schools deploying detection are both Vizzini, trapped in circular reasoning that goes nowhere.

    Zoom out. The poison kills you regardless of which glass it's in. Using AI to skip the thinking robs you of the learning whether caught or not. The detector is the least of your problems if you can't actually do the thing you're supposed to be learning.

    Show Your Process

    Work in drafts. Build the version history. The question that matters isn't "did AI write this?" but "can you defend this?"

    Here's what that actually looks like. Say you're doing a TOK exhibition. You can explain that you initially considered using your phone as an object to explore surveillance, but realized that was too obvious and didn't connect well to your prompt. You pivoted to your Spotify Wrapped playlist because it raised better questions about algorithmic knowledge. Then when you started writing, you realized you were focusing too much on what Spotify knows about you and not enough on whether that counts as knowledge at all. That's when you restructured the whole thing around the reliability question.

    That's complexity. That's evidence of thinking. AI can generate a polished commentary about Spotify and knowledge. It can't generate the trail of abandoned ideas and pivots that got you there.

    If you can walk someone through how your thinking evolved (where you got stuck, what changed your mind, why you made the choices you made), you're safe. If you can't, you've already lost something more important than marks.

    AI Can't Amplify What's Not There

    AI is a cognitive amplifier. It multiplies what you bring to it. Think of it like weightlifting. If you're a beginner who can lift 50 pounds, and you learn to use AI effectively, you might 10x your output to 500 pounds. That's a massive gain. An expert who can already lift 1,000 pounds without AI will still outperform you.

    When that expert learns AI and also gets 10x amplification, they're lifting 10,000 pounds. The gap widens if you let AI do the lifting for you instead of going to the gym yourself. Building your base capabilities matters because that's what gets multiplied. Working on your thinking now gives you a stronger foundation to amplify later.

    The Point Hasn't Changed

    The detection game is broken. Everyone playing it — students trying to evade, schools trying to catch — is trapped in the same unwinnable loop.

    The way out isn't better detection or better evasion. We confused the credential with the capability a long time ago, and detection is just the latest attempt to keep that confusion alive.

    The point was always the learning. Schools that build assessment around "can you defend this?" rather than "can we verify this?" don't need detectors. They need conversations.


    Sources

    1. Temple University Student Success Center — Turnitin accuracy audit
    2. OpenAI — Discontinued AI Classifier, July 2023
    3. Stanford University / Liang, Zou et al. — "GPT detectors are biased against non-native English writers," Patterns, July 2023
    4. Norwegian Directorate for Education (UDIR) — National directive on AI detector use
    5. MIT Sloan — "AI Detectors Don't Work" guidance
    6. Jack Cheng — "What Becomes Valuable When AI Makes Creative Work Easy," Every.to