AI in Practice

    Nowhere in My Job Description

    I recorded ten of my own lessons and ran discourse analysis on them. The real work wasn't the recording or the coding. It was catching the machine every time it lied.

    Charcoal illustration of the same man at two desks nine years apart — 2017, hand-coding paper transcripts under an amber lamp; 2026, the same colour-coded transcript glowing on a laptop screen.

    “What you're doing with these recordings is a larger version of the discourse analysis I did in my dissertation nine years ago.”

    A colleague said this to me in passing, the way you'd mention the weather. I thought to myself “hey that's neat, I did discourse analysis for my master's degree too!” because I had lessons to get through and the term was nearly over. It took weeks for the sentence to come back and sit down across from me.

    Here is what the recordings were. Over one term I recorded ten of my own lessons on top of teaching them. Three Theory of Knowledge, five English B at standard level, two at higher level. I ran the audio through ElevenLabs to get clean transcripts, pushed those through Gemini to code the discourse, and coded the questions by hand where the machine and I disagreed. Two hundred and fifty-five content questions in this round. I built a scorecard against six baselines I'd set for myself the term before, naming each one in advance so I couldn't move the goalposts after seeing the result. I wrote an audit log as I went, because I'd learned the hard way that you don't trust a number you can't trace back to the moment it came from. Every figure had to point at a sentence somebody actually said.

    I went in expecting good news. I'd spent the whole term working on one specific habit, and I wanted to see it move.

    It didn't move. The headline number said I'd got worse. I sat at my kitchen table that evening with the spreadsheet open and a slightly warm coke zero, reading the same row over and over, feeling the particular flatness of having worked at something and watched the data shrug. There's no clever way to describe that feeling. You just wanted the number to be better and it wasn't. I went to bed disappointed. It was only the next morning, brushing my teeth, that disappointment curdled into something more useful, which was the feeling that this couldn't be right.

    The habit was a verbal tic. Across the recordings I say “Cool. Cool, cool, cool” more than I'd like to admit, and the term before I'd seen the count and resolved to stop. This round the count came back higher. Two cools per file against one and a bit the term before. Worse, plainly worse.

    When I went looking for where the cools were hiding, I found that seventeen of the twenty-three raw instances came from a single lesson. That lesson was the one where I stood at the front of my class and read out my Round 1 results to my students, including, out loud, the sentence “I say Cool - cool, cool, cool too much.” The recorder was running. The new analysis dutifully counted every quotation. I had been measured saying the thing I was being measured for, while reading aloud the measurement that found it the first time. Strip the quotes and the real count was less than one ‘cool’ per lesson and only one cool cool cool in the whole set. The work was working. I just had to catch the machine catching me quoting myself.

    Two-panel chart. Left: cools per file appears to rise from 1.82 last term to 2.0 this term. Right: of all 23 raw instances across the term, 17 were quotations from a single lesson where I read my own results aloud; stripping the quotes leaves 6 — less than one per lesson.
    Sometimes averages don't tell the full story.

    The Theory of Knowledge number was uglier and slower to forgive. One lesson on the sixth of April had me answering seven of fifteen of my own questions. Just under half. For a teacher trying to get students talking, that made me feel like I'd fallen into a sort of “sage on the stage” format where I speak to entertain rather than to guide students. I almost let it stand, because actually I've fallen into that habit before and it's entirely believable.

    Then I went back to what the lesson actually was. It was a feedback conference, a format where you sit with the work and demonstrate how to read it. Four of those seven self-answers weren't me stepping on a student. They were modelling moves, rhetorical questions asked and answered aloud so the student could see my reasoning happen. Counting only the genuine ones gave four of eleven. Still high, still worth watching. Defensible, though, and behaving exactly as a feedback conference is supposed to behave. A rubric counts behaviours. Behaviour counts mean nothing until you know the context.

    The third one I caught only because I happened to know the size of the lesson. One audio file was thirty kilobytes. About thirty seconds of teaching, a fragment, something that had gone wrong in the recording. Gemini handed me back a twenty-five-minute lesson from it. Forty questions. False-uptake counts. A rephrase rate of thirty-eight percent. Timestamps that ran past the twenty-four-minute mark, confident and specific and entirely invented, because there was nothing past second thirty for any of it to be about. The model met a void and filled it with priors. None of it had happened. I pulled the row out of the dataset, and when I did, three of my six baselines flipped category. The fabricated lesson had been quietly suppressing real progress. Two of the six still went the wrong way after the strip, on sample mix and wait time, though wait time was still a healthy 4.6 seconds. The point is that without a human who knew what was actually in the room, I'd have been optimising my teaching against a lesson that never took place.

    Dumbbell chart of six self-set baselines before and after removing the fabricated lesson. Three verdicts flipped to hits — HL deep questions, questions per lesson, teacher talk. Sample mix and wait time stayed in the wrong direction.
    One fabricated row, pulled. Three of six verdicts flipped. The fake lesson had been suppressing real progress.

    That's the work. Not the recording, not the coding, the interrogation. Sitting with rows of data enough to ask where was I when I said that. What do you call the skill of taking a fluent-looking output and distrusting it? I picked that question up from Dan Klein, and as far as I can tell nobody has named the skill yet. It's most of the job now.

    The point is not that AI made analysing my teaching more efficient but that AI has given me the power to redefine my professional role.

    I'm a teacher. My job description is about what you'd expect. Nowhere in it is “conduct discourse analysis on over forty hours of teaching and use it to optimise lesson transitions and questioning style.”

    Which brings me back to my colleague, and the sentence that took weeks to land. Nine years ago I did discourse analysis for a dissertation on implementing the IB learner profile in a bilingual school. The same craft: transcribe talk, code it, look for the patterns underneath what people think they're saying. I got good at it through 10s of hours of practice and feedback from my advisor. And that's really important to be honest about. The skill didn't appear this year. It was already in me, sitting unused for the better part of a decade, because the version of it that fits a working teacher's life — forty-plus hours of your own classroom, one term, alongside the actual job — used to require time and tools and a research budget I was never going to have. The discipline never changed. The bandwidth did. What AI removed wasn't the gap in my ability. It was the gatekeeping between a thing I already knew how to do and the work itself.

    That's why I won't pretend this is a universal key. It scales exactly as far as your own training and skill does. The colleague who named it could only name it because she'd seen it before, in a context where she'd done it before. If there's no earlier version of this in you, AI doesn't conjure one. It unlocks what's already there and stranded.

    The boundaries between roles are blurring, and they will keep blurring. I'm less interested in that as a forecast than as a question I can now answer about myself, which is what I'm actually allowed to do for a living.

    I'm a teacher who does education research now. I wouldn't mind to try my hand at a few other roles, too.