AI in Practice

    The Compound Anthropic Forgot to Split

    Anthropic drew a product around the shape of my week — and a bright line under the part that was ever actually mine

    Charcoal illustration of a teacher's satchel split down the middle: the left half drained grey, fed into a machine claw beneath smokestacks; the right half warm cognac leather under a brass lamp, with a fountain pen resting on marked pages.

    In a talk to the hospitality industry, Rory Sutherland has a line about electric cars. Range is a matter of engineering and physics. Anxiety is a matter of psychology. Two unrelated problems hide inside one complaint, and if you hear "range anxiety" as a single thing, you build a bigger battery when the driver needed a better map.

    "Teacher workload" is that kind of word. It sounds like one job. It is at least two.

    Yesterday Anthropic launched Claude for Teachers: a free year of premium Claude for US K-12 educators, a standards-aware lesson planner, a skills library built with learning scientists, scheduled tasks that run while you sleep. If you teach anywhere else, as I do, I'll save you the scroll: there is no international version and no timeline for one. Keep reading anyway, because that turns out to matter less than it seems. I have been building versions of nearly all of it, by hand, for eighteen months. My first honest reaction was not a threat but recognition. Someone at Anthropic had drawn a product around the shape of my week — a real teaching week.

    My second reaction took longer, and it is the reason for this piece. That word "workload" smuggles in two things. There is the admin: the highlighting, the retyping, the reformatting, the standards cross-referencing, all of which I would happily never touch again. There is the judgment: knowing this class, hearing what a student half-understood, deciding what a comment should actually do to them. Those two live under one word, and any product built to reduce "workload" will reach for both unless something stops it.

    The thing that stops it is the number you decide to watch. Anthropic's design metric, everywhere you look in the launch, is time. The flagship demo is measured in a phrase: review today's exit tickets, adjust tomorrow's plan, done by 4pm every school day. Time saved is a clean, honest, demoable number. It is also the number that quietly decides what gets built. Optimise a teaching tool for time and the logic points at the teacher, because the teacher is the slow part. Optimise the same model for ownership and it points at the admin instead, because the admin is the part nobody should have to own. Same model, opposite products, and the metric is what chooses between them.

    Reading the feature page was like reading a description of my own house written by someone who has never been inside: the rooms all there, the furniture slightly wrong, a wall or two out of place, but still unmistakably where I live. Here is the overlap.

    Advertised in Claude for TeachersWhat I already run by handVerdict
    Lesson planning against standards (the Learning Commons "connector", the product's data plug-in, carrying every US state's standards plus two major curricula)Pasting IB subject guides plus class context into promptsPartial. The workflow, yes; the standards dataset, no
    Differentiation (3-tier materials, reading level, ELL)LessonImprover: bilingual EN/中文 Tier 2/3 vocab, EAL scaffolding, IB-awareSame job; mine speaks two languages
    Exit-ticket review → adapt tomorrow ("4pm daily")Scheduled exit-ticket jobs on my own infrastructureSame pattern, my hardware
    Assessment generation (standards-tied quizzes, keys)ExamGenerator: IB Paper 2, 11 question types, student/teacher packsSame job; mine is built to the IB exam format
    Grading / feedbackTeachingAssistant plus my speak-to-margins toolNot shipped; this is where most of my building went
    Teaching skills library ("built with learning scientists")My entire SKILL.md librarySame primitive, productised
    Data-driven instruction ("hand Claude a folder")Results-calibration analysis, feedback data workSame
    Lesson transcripts (TeachFX connector)Two terms of recording my own lessons; transcripts fed feedback on my teachingSame pattern; mine audits the teacher
    Polished materials (Claude Design)Documents skill plus my house brand pipelineSame

    Seven full overlaps, one partial, one where the product does not yet do the thing at all.

    One piece of context the launch coverage keeps missing: Anthropic did not invent this category. MagicSchool, Diffit and their cousins have been selling standards-aligned planning and differentiation to teachers for two years. MagicSchool is on the launch page as a partner, and per Anthropic's own case study it has already moved its product onto Claude under the hood. Consider that. The company that built the category now runs on the platform that just entered it. If you want to know what platforms eventually do to the products built on top of them, you don't need my predictions; you can watch it happen on this launch page.

    Two rows deserve a closer look, because they are where the split shows cleanest.

    The row the product barely ships, and I run the most

    Start with feedback. Picture the actual thing. A stack of twenty essays and by number fifteen I am usually crawling over the finish line. The same comma splice appears in fifteen of the twenty texts, and my patience for explaining it kindly is teetering on the edge of a chasm of grammar errors. The feedback gets shorter, and the fifteenth student pays for a fatigue that has nothing to do with their essay.

    The tool I built for this, which I call teachspeak, does something much narrower than "grade my essays." I read the script and talk while I read. My speech becomes a table of comments, and a Word macro drops each one into the right margin connected to the right text. It removes the highlighting, the typing, the rephrasing of a blunt thought into student-facing English. It does not, I want to be precise, speed things up a whole lot. What it changes is that I finish the stack without crawling out of school. The judgment stayed mine the whole way. Only the typing left.

    The most revealing part is what the machine does to tone. It always cuts the snark. Most of the time it is right, because the snark was fatigue. Sometimes the sharp comment was doing honest work, because a student who coasts needs to feel a little friction, and I put it back in by hand. The machine filters the fatigue artifacts. I curate the edge back in. The popular fear is that AI sands teacher feedback down into "game changer," "synergies," "delves," and other corporate lingo. In practice it hands me a cleaner surface to decide the tone on.

    Here is what that looks like from a real stack, a set of TOK exhibition drafts back in May. What I said, somewhere around the sixth script: "For the final object, the Apple Watch, there's really not much to be said here. You could say the exact same thing about any other heart rate monitor. Don't be that mean writing it to the student, but for the sake of making my point clear, that's the case. It's just not got any substance to it." What landed in the margin: "This object lacks specificity—the same point applies to any PPG heart rate monitor. Consider finding an object with unique, deeper meaning to explore." The "no substance" is gone, and it should be. The heart-rate-monitor point stays, because the student needs to know why the object fails, not just that it does. And the sorting is right there on the tape: I decided out loud, mid-comment, which half of what I felt belonged on the student's page. That call is the entire job, and it never left my desk.

    A real teachspeak session: a dark transcript strip quoting what was said aloud about the Apple Watch object, above a Word-style document with the student's paragraph highlighted and two softened margin comments from Mr O'Neil.

    That points at a test you can run on any AI feedback workflow. Talk about a class's writing three days later, without opening the file. If you know that half of them still can't punctuate dialogue and two of them have made progress in their weakest grammar points, you owned the feedback. Paste the whole stack into a chatbot and you fail that test by the next morning, because you were never where the knowing happens. The distinction between being a writer and being an editor is much bigger than it looks. Paste-in feedback makes you the editor of comments you didn't write. Teachspeak keeps you the writer and makes the machine the typist. Both save time. Only one leaves you the author of your own judgment.

    The row Anthropic put on the poster

    The second row is the demo everyone will screenshot, and it is where the split cuts deepest. Claude reviews today's exit tickets at 4pm and adjusts tomorrow's plan, every school day, automatically. It is a genuinely good demo. It is also a small machine for skipping the most valuable fifteen minutes of my afternoon.

    Charcoal illustration under a clock reading 4pm: a grey robotic hand fans out three paper slips while a warm human hand under lamplight reaches to take one of them.

    There is an old story about doormen. A hotel decides the doorman is inefficient, since a door can open itself, and removes him, then discovers the doorman was also hailing taxis, spotting the regular who had lost his key, turning away the man who didn't belong. The open door was the line item. Everything else was invisible value that never appeared in the savings column.

    Reading thirty exit tickets is my open door. The adjusted plan is the artifact. What I actually walk away with is thirty small windows into what landed and what didn't, and by ticket twenty I can feel tomorrow's lesson rearranging itself in my head. Sutherland has a companion point about university essays: they are close to worthless as essays, and what was valuable was the work you had to do to write them. Exit tickets run the same trade. Hand the reading to a scheduled job and you keep the adjusted plan while losing the reason it was the right adjustment. The byproduct outvalued the artifact the whole time.

    The demo is one verb away from good. It fails as shipped because Claude reads and then decides. It passes the moment Claude reads and then surfaces three things it noticed, laid on my desk at 4pm — and I decide what tomorrow does with them. A few words of difference in the spec. The whole teacher, kept or removed, hangs on those words.

    What actually dies

    Honesty demands I say which of my tools this product kills, because a piece like this slides too easily into a man insisting his job is special. Some of it should die, and will.

    Exam generation goes first. I built a generator for IB English B papers, and I am not precious about it. Writing good exam questions is real craft, the kind teachers two years in don't reliably manage, but a standards-tied generator will get close enough to make the hobby version pointless. My speak-to-margins trick may already be native to the product now, and that is genuinely worth me checking rather than defending. The differentiation prompts, the lesson scaffolds, the polished handouts: those are packaging over primitives anyone can run, and packaging is precisely what a platform does better than a person.

    The row I nearly scrolled past is the one where the connector had a familiar shape. TeachFX turns lesson audio into transcripts, and I spent two terms doing exactly that to myself — recording my own lessons and running the transcripts through the same kind of pipeline I use for marking. It is how I know that across two months of lessons I said "Cool. Cool, cool cool" twenty-four times, and how I ended up one afternoon staring at a spreadsheet of 702 coded questions from my own classroom. The product points the transcript at the class, but the version that changed my teaching pointed it at me.

    What survives is narrower and harder to copy, and it isn't the workflows at all. It is the judgment those workflows were carrying: which comment lands on which student, which snark is honest, which fifteen minutes I refuse to schedule away.

    I should be clear about my own position in this. I run most of my setup on Claude, by choice, because I think Anthropic is building the right things. I am even planning a small local-model node at home for the dumb, high-volume jobs, precisely because the consumption logic in this post applies to me too. The more of the simple work a platform swallows, the more I want to own the boring half myself.

    I said at the top that there is no international version, and there is a second fact to sit beside it: the free year is US K-12 only and the price after it is undisclosed, which makes it a window rather than a gift. The honest message across the border turns out to be a simple one: you can't have the SKU yet, but you can have the practice today.

    The accountability version

    So here is the promise, on the record so I can't quietly drop it. I am going to write up my tools one at a time, the feedback pipeline and the exam generator and the exit-ticket loop and the differentiation stack, and for each one I'll say plainly whether Claude for Teachers ate it, matched it, or missed it. No grading on a curve for the home team. The feedback pipeline goes first, and this time next week you can check whether I kept the promise. Partly because the write-ups will be useful, but mostly because claiming there is a why under each tool and then showing it is the only way to be sure the why was ever there, and not just habit dressed up as judgment.

    Anthropic built a product around what I was already doing, and in the same motion drew a bright line under the part that was ever actually mine. The workflow was always the easy half to copy. The judgment it was carrying is the half that walks out of the building when I do. Eighteen months of doing all of it by hand bought me the one thing the free year doesn't come with: when the SKU finally crosses the border, I'll know exactly which boxes to help it carry into my classroom, and which one stays on my desk.