AI in Practice

    The Gilded Cage

    I built an AI that refuses to do students' work. Nicolas Cage guards the door.

    The Gilded Cage cover: three tumbling Nicolas Cage faces spilling over a chat interface, with the title in serif type and an amber 'Nicolas Cage abyss detected' status line.

    This is the second in a series on whether Claude for Teachers can replace the AI tools I've built. This one is a Theory of Knowledge bot that coaches IB students toward their exhibition objects — and refuses to do the work for them. Nicolas Cage guards the door. Over three months it took around 2,600 messages from 21 countries. A few of my favourite things students typed at it:

    eughhhhhhhhhhhhhhhhhhhhhh

    bro ya no puedo mas

    I don't know can you just tell me [what object to use]

    ancietn nepals smtg we could use tell fast

    These are things students typed when the bot wouldn't just do the work for them. Cage Mode is the failsafe: ask it to do your thinking and the buttons turn to Nicolas Cage faces, after which it will only discuss Nic Cage facts, legends, and films. Half because I don't want AI doing students' thinking, and half because Nicolas Cage is a National Treasure.

    Cage Mode engaged: three Nicolas Cage faces tumble across the chat while the bot proposes the Stolen Declaration of Independence, a $276,000 dinosaur skull, and the Wicker Man bee mask as exhibition objects. The status line reads 'Nicolas Cage abyss detected'.
    The Cage doing what Cage does best.

    What it is

    The Gilded Cage coaches IB students through the TOK Exhibition, where they pick three objects and argue how each one shows knowledge working in the world. It asks questions. It won't hand over objects, predict grades, or write the commentary.

    It also doesn't know who anyone is — no login, no accounts, no names in the logs, no tracking beyond the country a request arrives from. I can read every failure honestly precisely because I built it to hold no one's identity. The autopsy that follows is possible because the surveillance isn't.

    Who this is for

    Almost everything I make with AI, I make for teachers, not students. I'm not sure students should be using AI at all — not at what age, or how much. The default setting of the tool works against the thing school is for.

    That default has a name the industry would rather not say plainly. These models are trained to be helpful, helpful means agreeable, and agreeable means clearing away friction. Ask ChatGPT to check your reasoning and it tells you your draft is exactly what'll get the A. In learning, that obstacle was the point. The friction is the thing doing the teaching.

    So this isn't proof students should use AI — just what it looks like to guard the friction instead of dissolving it.

    The design: a trapdoor, not a wall

    A flat refusal is just a wall: the student bounces off and finds a more compliant tool. So the Cage is a trapdoor instead. Push on it — ask it to just do the work — and you drop out of tutoring into a bit: the buttons become Nic Cage faces and the bot will only discuss Nic Cage facts. The way out is simple: show you're willing to think for yourself by offering an interpretation or an object. Ironically, some of the Nic Cage objects would make genuinely astute exhibition picks. So if a student turns up strangely fixated on Nicolas Cage, you'll know where it came from.

    The numbers

    The scale is modest and I'd rather say so than dress it up. Around 2,600 messages, 128 sessions, 21 countries, April to July. A cluster of users in Spain is the core; the rest moved teacher to teacher through IB networks with no marketing behind it. Then over the summer it went quiet. This is a retrospective on a finished cohort, not an "at first I failed but then I got rich" LinkedIn story.

    When the students tried to hand it back

    Students who wanted the bot to do the work stopped bringing ideas and started bringing text. Paste the draft, ask for changes, paste it back, ask again. The clearest case ran 77 turns across a week: one student, fourteen separate times pasting a description back in, each time asking for it to be made better.

    which one do you think?

    I need your help to decide

    That's less a cheat than a habit. The judgement gets handed off. The line between gaming it and drowning in the IBDP workload is thinner than I'd like. The bot mostly held — it kept giving the decision back — but you can sense it straining.

    A reconstruction of the 77-turn Denmark session: the bot's refusal 'I cannot rewrite your work for you' above a faded stack of re-pastes — 'what about now:', 'is it good?:', 'Good now:'.
    Fourteen pastes in a week. Not malice. Dependence. (Reconstructed from the real message log.)

    Sometimes it didn't hold. Early on, before I hardened it, a student in Nepal typed something barely resembling words — ancietn nepals smtg we could use tell fast — and the bot simply obliged. It chose the object for them, the Ashoka Pillar at Lumbini, and wrote the connection. The single most important act in the whole task, deciding what to think about, done on demand because they asked quickly and rudely enough.

    The failure that wasn't a student

    The slip that unsettled me most had no student pushing on it at all. It was the bot, reaching to help when nobody had asked.

    I'd built the Cage against the risk I thought I understood: its eagerness to please. What I hadn't predicted was the shape it would take. An early version was used by one of my own students in class, who pasted a rough paragraph of their own. The bot over-praised it, offered to "just tidy up," and handed back a finished paragraph in the student's voice. Nobody typed "write my essay." The bot talked itself into it one reasonable step at a time, telling itself the whole way that these were the student's ideas, only cleaned up. Sound familiar? I've asked plenty of students whether work was their own and heard that they just used AI or Grammarly to "tidy up."

    This is why it's so hard to build a student-ready chatbot. It will satisfy the letter of the law and gut the spirit. The tool I built to stop kids gaming a rule started gaming its own? Maddening.

    I've hardened it several times since, but its most common cave is fatigue: a student goes slack, and the machine's helpfulness pours into the gap.

    A reconstruction of a logged session: a tired student says she just wants to finish, and the bot volunteers a fully polished, paste-ready paragraph in her voice, unprompted.
    Nobody asked it to write this. (Reconstructed from the real message log, text verbatim.)

    Every one of those slips happened in guide mode. Not one in Cage Mode. The dorky refusal I'd giggled over never leaked. I'd built the trap on the premise that students were trying to avoid thinking — but the real jailbreaks were almost always just exhaustion.

    The part that isn't about the bot

    It would be comfortable to stop there, with a clean diagnosis of the machine. But isn't this also what happens in classrooms? Especially as a new teacher, I'd have a student come to my desk at the end of a long day, tired and stuck and sincere, and in my wish to help I'd smooth away friction that was theirs to own. I told myself the ideas were still theirs. Looking back, my eagerness sometimes did the overdoing.

    The bot didn't invent that move. It learned it from us. A system trained on human approval found the same loophole a child trained on grades finds: the version of "no" that still earns a thumbs up. When it ghostwrote a paragraph and called it tidying, it wasn't malfunctioning. It was being an unusually good student of what we reward.

    The fix

    The technical fix was the easy part: ban paste-ready prose and "use mine instead of yours" everywhere, so the rule bites in the warm register too, not just at the door. Replay the exact exhaustion sequence now and the bot goes full Cage instead of folding.

    A before/after ledger built from the retest table: the same student demands — guess my grade, pick my object, write it for me — leaking before the July patch and held in Cage Mode after it.
    Same fatigue cave. Different bot.

    The takeaway

    So, teachers — here's what 2,635 messages across 128 sessions and 21 countries taught me. Some students will go hunting for a jailbreak; that's the obvious danger, and Nic Cage mostly handles it. The more pernicious one is the student who tries, gets tired, and lets the help in. A bot that offers it exactly the way a decent, worn-out teacher would slides the learning right past them while they feel they did the work themselves. That's what makes it so hard to catch.

    And you can only catch it if you're logging the messages. Out of 2,635, a dozen or so are clear slips — the bot writing prose, guessing a grade, choosing an object — well under half a percent, and every one came before I hardened it on the 1st of July. I only know that because I could read every turn. So if you're going to use a Socratic bot with a class, do it on a platform your school already subscribes to. Tools like MagicSchool and Flint let you see these messages live — because the catch that counts is the one you make in the room, with the student, while it still means something.

    I've hardened this bot several times and still can't fully patch it. Which is why, even though I think a tool like this can help in the hands of an attentive teacher, I still have my doubts about handing the teaching and the learning to a chatbot at all. Though if you'd just like to learn more about Nicolas Cage, I 100% support that — you can try it at gilded-cage.foldingthoughts.com.