Teaching & AI

    The Decathlon

    How AI helped me rebuild a lesson on disadvantage that finally landed.

    Article Preview
    0:001:06
    The Disadvantage Decathlon — an equalizer with 13 capability sliders representing Wolff's framework for measuring disadvantage

    This post is about how I built the Disadvantage Decathlon — an interactive web tool for teaching Jonathan Wolff's framework on disadvantage. Nine real communities, thirteen sliders. You can try it yourself here.

    About five years ago I was listening to Philosophy Bites. It is a podcast where Nigel Warburton interviews philosophers in 15-to-20-minute episodes that are dense. More than dense, they are cerebral. I feel like they are the podcast equivalent of flying a fighter jet. Like you strap in and for the next 15 minutes you are listening with precision. Probably a bit less intense physically though. Anyway the episode was with Jonathan Wolff, discussing his book Disadvantage. In it, he explains the decathlon model: his framework for understanding disadvantage not as one thing, but as a composite of many capabilities. He uses the metaphor of a decathlon to say that no matter what you weight heavier (health, access to education, bodily autonomy), the same groups end up at the bottom of the ranking. The same groups at the top.

    I found this so compelling that I bought the book and read it. The framework stuck with me for years because it does something rare: it takes a concept most people intuitively understand — disadvantage — and reveals the measurement problem hiding inside it using a brilliantly simple analogy. When it comes to understanding and learning, I think metaphors and analogies are incredibly useful and this one struck me as especially brilliant.

    One of the things I enjoy most about teaching is taking ideas like this (grad-level, nuanced, not originally intended for 16-year-olds) and figuring out how to make them land. This is the challenge that gets me out of bed. Not simplifying the idea, but finding the right form for it.

    The first time I taught it, I used too much of the book. Longer extracts. Dense philosophical prose. The framework was intact, which felt like a win until I noticed students checking out by mid-lesson. The next year I trimmed the reading, shortened the excerpts. It helped, though the core problem remained: the decathlon model is fundamentally about interaction between categories and empathizing with people vastly different to you. Explaining this on paper is like explaining music by describing the notes.

    This year, I rebuilt the lesson with AI.

    The Diagnosis

    The first thing I asked AI to do was tell me something I already half-knew:

    claude code
    The decathlon analogy is the best part and it's buried. The idea that
    disadvantage clusters — that no matter how you weight the categories,
    the same groups end up at the bottom — is genuinely surprising and
    provocative. But it appeared on slide 12 of 17, after students had
    mostly checked out.
    
    No student-facing activity that makes them DO the thinking. They read,
    discuss, define terms. But they never try to MEASURE disadvantage
    themselves and discover why it's hard. The lesson tells them it's hard
    instead of letting them fail at it.

    That last line hit. The lesson tells them it's hard instead of letting them fail at it. This is the kind of thing you feel as a teacher: the best learning happens when students encounter a problem directly, not when you describe the problem. I'd been describing the problem for three years.

    The Collaboration

    What happened next is what I think is worth sharing about working with AI. It wasn't a case of "generate me a lesson." It was a conversation, and the conversation required me to hold the line on things the AI couldn't know.

    When it suggested ranking activities, I flagged the ethical risk:

    claude code
    We gotta be pretty careful with how we frame this because if it feels
    like we're literally doing a form of the poverty Olympics, it needs to
    be framed as a policy-maker exercise. And we have to be careful this
    doesn't become a political science taster lesson. It needs to still
    focus on this idea that this is about scope of human sciences —
    we're still on scope.

    Three concerns in one breath: ethical sensitivity, pedagogical substance, curricular alignment. The AI proposed activities. I kept them grounded.

    When it suggested dropping the "obvious" categories of disadvantage, I overrode it. Those categories matter precisely because students think they're obvious, and then discover they're contested.

    When I said I wanted real community data, not hypothetical profiles, I launched four AI research models in parallel:

    claude code
    ⟩ claude-researcher  — "Research disadvantage communities"
    ⟩ gemini-researcher  — "Research disadvantage communities"
    ⟩ gpt-researcher     — "Research disadvantage communities"
    ⟩ grok-researcher    — "Research disadvantage communities"

    Four models searched simultaneously, each bringing different sources. The prompt I gave them specified the academic framework, the three-tier design I wanted (communities that will likely rank bottom, likely rank top, and communities that drift depending on what you weight), ethical constraints ("respectful and representative, not sensationalized"), and some must-includes: the Sundarbans honey collectors in Bangladesh, who risk death by tiger for $70 a season and the Bedouin of the Israeli Negev. These communities are directly from Wolff's book, and it's the kind of detail that makes abstract philosophy visceral.

    The four models didn't all agree. Their synthesis looked like this:

    multi-model synthesis
    Drift Candidate        | Claude | GPT | Grok | Gemini | Consensus
    Appalachian            |  Yes   |  —  |  —   |  Yes   |   2/4
    Aboriginal Australians |   —    |  —  | Yes  |  Yes   |   2/4
    Sami                   |   —    | Yes | Yes  |   —    |   2/4
    Inuit                  |  Yes   |  —  |  —   |   —    |   1/4

    Aboriginal Australians emerged through multi-model consensus, a stronger pick than the Inuit originally suggested, because the drift is more dramatic. Weight cultural sovereignty: they rise. Weight health outcomes: they collapse. That tension is what I wanted the tool to have to teach.

    The Moment It Worked

    The tool is a web app. Nine real-world communities. Thirteen sliders for Wolff's capability categories. Drag a slider and watch the rankings reorder in real time.

    The AI had labelled each community as "Always Top," "Always Bottom," or "Drifter." I told it to remove the labels:

    claude code
    I don't think we should write right there that these groups are always
    on top, always on bottom, drifters, etc. Let people play and discover
    that. You can set up the system so they're NOT top. I did it where I
    turned everything off except for Affiliation and the Aboriginals and
    Bedouin come up at the top. How interesting.

    The labels would have announced the answer before students found it. The whole point of the decathlon model is that the answer depends on what you value. Spoiling that with a label defeats the purpose.

    What Changed

    The tool makes the lesson more empathic. Each community profile gives students a window into lives genuinely unlike their own: honey collectors navigating tiger-infested mangroves, Gulf state migrant workers earning wages they can't spend freely, Appalachian families with full citizenship and drastically worse health outcomes than the rest of the country. The abstraction of "capability categories" becomes concrete when attached to real people in real places.

    Once I started using it, a question nagged at me: how were these scores determined? If a student clicked on the Rohingya and saw a 1 for Bodily Health, they should be able to ask why — and get an answer. So each community now links to the research behind its scoring. UNHCR data for the Rohingya. AIHW life expectancy studies for Aboriginal Australians. Wolff's own case studies for the Sundarbans and the Bedouin. The scores are informed estimates, not precise measurements, and the tool says so. But the reasoning is visible. That felt important to me. You can't teach students to interrogate knowledge claims while hiding your own.

    This lesson has always been one I look forward to teaching. For three years it never quite landed the way I wanted. The idea was always there — Wolff's framework is genuinely powerful. The packaging wasn't. What changed wasn't the idea. What changed was the form. And AI was what allowed me to change that form.

    If you teach and you're experimenting with AI for things like this, I'd like to hear about it. The design decisions that make a tool work in a classroom only come from teachers who know where their lessons fail. That knowledge is worth sharing.