Course Insights

    The Illusion of Structure

    Why do half the TOK essays I mark look right and read empty?

    An earnest student stands alone on a stage in an empty exam hall, fully kitted out in five well-meaning but empty signals: a DOUBT sticker on the chest, a HUMAN SCIENCES party hat, a DOUBT DRIVES BREAKTHROUGHS sash, an open dictionary with the wrong-level definition, and a clipboard of ungraded bullet points.

    Teach TOK for long enough and you start to see the same mistakes repeating. The essays that fail almost always look right and that's why it's quite innocuous. The trouble compounds because the most common feedback teachers see on essays they buy back from IB feel like coded language you need to decipher. “The argument is too descriptive”, “the argument is not about knowledge and/or it's not about knowledge in the given AoK.” These sorts of comments are essentially why you can write an essay that reads well and seems to show some strong logical argumentation and yet scores 2–4/10.

    The students who write essays like this have learned the form of an essay and trust the form. Thesis up front. Two AOKs named. Key term defined. Conclusion that circles back. Everything the form asks for, delivered. What the rubric is asking for is something else, and most essay struggle to demonstrate it. That can be seen clearly in the global average score on the essay.

    What follows are the five most common patterns I saw, with one real anonymized paragraph from a student essay each. If you teach, these are the mistakes worth naming for your students. If you write, they're the same five traps that show up in any kind of analysis.

    1. The keyword sticker

    The title's key term gets used as a label on cases that are about something else.

    Picasso started to doubt whether this realism depiction was the only way to make and represent art. In Les Demoiselles d'Avignon, he combined many different viewpoints in a single art piece, using angular forms and taking inspiration from African masks and art traditions that were not Western. Picasso's doubt of the norm led to the creation of Cubism, an art movement that redefined how artists represented shape, space, and form.

    The prescribed title is about whether doubt is central to the pursuit of knowledge. This paragraph treats a famous artistic move—Picasso combining viewpoints, working from African masks, breaking realism—as evidence that he doubted. But making different art is not the same as doubting. Swap doubt for experimented, innovated, or rebelled against and the paragraph reads exactly the same. The student names what Picasso did. They never engage with what he questioned, what evidence shows the doubt operated, or how doubt-as-mechanism differs from inspiration-as-mechanism. The keyword is decorative. It sits on top of the case rather than picking it apart.

    Diagnostic: swap the title's key term for a cognate. If the paragraph still works, the term isn't doing analytical work.

    2. The wrong definition

    You define the key term, but at the wrong level.

    Doubt refers to the uncertainty in one's ability to complete a task. Personally, I encountered doubt at many points in school, but especially during the transition from middle school to high school. Examinations became increasingly important, raising my stress levels and consequently exacerbating my existing doubt.

    The prompt is asking about epistemic doubt, the kind that drives revision of knowledge claims. This essay defines doubt as task uncertainty, the kind that makes you nervous before an exam. The rest of the essay then goes on to describe Pasteur and Kahneman, who clearly weren't doing this type of task-doubt. The intro and the body are answering different questions.

    Diagnostic: does your definition operate at the level the prompt is asking about? Does it distinguish your key term from its closest cognate?

    3. Wearing an AOK Party Hat

    The example sits inside the claimed area of knowledge topically, but the paragraph engages no theorist, theory, or method from that discipline.

    Platforms such as Facebook and Twitter used algorithms that only exposed users to content they had previously shown interest in. This selective exposure to information reduced users' capacity to doubt what they saw and led to the spread of misinformation, including the “Pizzagate” conspiracy theory.

    This is filed under human sciences. It names no human-sciences researcher, no theory of epistemic communities, no methodology. A bright sixteen-year-old who's never studied any human sciences could write this paragraph from what they already know. Topic-in-AOK without frame-in-AOK.

    Diagnostic: delete the AOK name from the paragraph. Does it still clearly belong there via named concepts, methods, or researchers?

    4. Assert without demonstrate

    You label a case as illustrating the key term without showing what the knower actually did.

    Two centuries ago Lobachevsky and Riemann started to doubt the fifth postulate of Euclid—the assumption that parallel lines never meet. Their doubtfulness led to new mathematical knowledge called non-Euclidean geometry. Therefore, doubt is a driving force behind major mathematical breakthroughs, which would not have happened without uncertainty and skepticism.

    The student tells us doubt drove the discovery. The student does not tell us what Lobachevsky or Riemann actually did. They could just as well have been driven by curiosity, or by the aesthetic frustration of an axiom that didn't pull its weight. Doubt is a label assigned to their work, not a mechanism shown operating.

    Compare with this, from the same prompt:

    In the 1840s, physician Ignaz Semmelweis questioned the prevailing “miasma” theory of disease and accepted practices in contemporary clinics. Semmelweis discovered that the death rates from post-childbirth fever were many times higher among women attended to by physicians than among women attended to by midwives. His doubt about his community's practice led him to hypothesize that the physicians were transferring “cadaverous particles” from postmortems to their patients.

    Semmelweis did something specific. Noticed an anomaly. Formed a hypothesis. The doubt is doing visible work. The Lobachevsky paragraph asserts doubt; the Semmelweis paragraph demonstrates it. The difference is small in word count and large in score.

    Diagnostic: for each case in your essay, can you name what the knower actually did?

    5. The list, not the argument

    You stack perspectives with caveats but never weigh them against each other.

    Survivor testimonies provide first-hand perspectives, but are shaped by trauma, fear, and the passage of time. Media coverage can be immediate and vivid, yet it may focus on dramatic or selective aspects, reflecting editorial choices. Official records, while seemingly authoritative, may reflect political interests or incomplete knowledge available at the time.

    Three perspectives, three caveats, no adjudication. The paragraph reports that the perspectives differ. It never says which one carries more weight under what conditions, or what criterion you'd use to decide. Listing perspectives is not evaluating them. Evaluation requires a defended criterion.

    Diagnostic: when you name competing perspectives, do you offer a principled criterion for weighing them, and defend it, or do you just announce that they differ?

    The single move underneath all five

    If you reread the five examples, the same failure shows up in different costumes. Each one uses the prompt's key term as a label for something already happening, instead of as a lens that picks out which cases count and how to read them. Keyword sticker is the shape at sentence level. Wrong definition is the shape at intro level. AOK costume is the shape at paragraph level. Assert-without-demonstrate is the shape inside a single case. Listing-not-evaluating is the shape across several cases.

    The fix is the same in all five places. Take the prompt's key term, define it operationally, and ask of every paragraph: is this paragraph showing the term operating as a mechanism, or just sitting on top of an example like a name tag?

    When I mark a TOK essay, I'm watching what the keywords or key phrases do from paragraph to paragraph. When they do the same thing each time, the conclusion rarely moves the score off four-to-five.