A recent MIT Technology Review article makes a strong claim “Don’t be fooled - LLMs don’t reason”. I think the argument is much better than the headline, and it deserves a careful reply.
TIP: Checkout the SKILL.md at the end of this post…

Its author, Thore Graepel, was a core member of the DeepMind team that built AlphaGo, and he opens with a moment he watched from the inside - move 37 in game two against Lee Sedol in 2016. The move looked so strange that some commentators assumed it was a glitch. AlphaGo’s policy network, the part trained to guess what a strong human would play, rated it at roughly one in 10,000. What actually chose it was AlphaGo’s search - an explicit tree of possible futures, each position weighed by the networks’ judgements. Graepel’s point is that this wasn’t a flash of machine intuition. It was reasoning: fast, learned hunches tested against a slow, deliberate exploration of consequences.
Today’s LLMs, he argues, only have the first half of that machine. They predict the next token, and even chain of thought is just more of the same process run for longer, with no separate machinery keeping track of hypotheses, evidence and uncertainty.
He’s right that this matters. But I think his argument rests on two assumptions that are becoming harder to defend. The first is that a large model is essentially a very sophisticated System 1 (his term) - fluent pattern completion without any deeper reasoning machinery. The second is more subtle: that the model is the system. I think the second assumption is likely the more important of the two.
Where we agree
Before disagreeing, it’s worth being clear about how much of Graepel’s argument I accept. In medicine, engineering and science it matters not only what a system concludes but how - when something goes wrong, we need to know whether the reasoning failed, the evidence was bad or an assumption was wrong. Beliefs should change because evidence changed them, not just because a fluent paragraph made them sound more convincing. And I agree that making intuition bigger doesn’t, by itself, make it more deliberative.
So the question isn’t whether trustworthy reasoning needs explicit, auditable epistemic machinery. For the work he cares about, it clearly does. The question is where that machinery should live, and whether we already have what we need to build it.
Next-token prediction describes the interface, not the computation
“The model only predicts the next token” is true, but it’s also really just unhelpful. It describes the training objective and the interface. But it tells us very little, if anything, about the internal computation the model has learned in order to meet that objective - in the same way that “a chess engine flips bits in memory” is true and says nothing useful about how it plays chess.
If the best way to predict the next token in some domain is to track objects, relationships, hidden states or competing possibilities, then next-token training puts real pressure on a model to learn machinery that does exactly that. And over the last few years we’ve found that it does. A transformer trained only on Othello move sequences developed an internal board representation that researchers could edit to change its moves (Li et al.), and later work showed that representation is largely linear. A model trained on chess notation learned the board and even latent player skill. Llama models turned out to organise cities, landmarks and historical events along coherent internal maps of space and time.
Of course, none of this proves an LLM holds one complete “world model”, and I actually think that phrase often points us in the wrong direction altogether. The more useful picture is many smaller, reusable structures that I’ve been calling latent models - internal scaffolds that represent relevant states and relations, support prediction, and respond coherently when you intervene on them. The important word here is usable. A representation that merely correlates with the answer is interesting. One the a model actually computes through is much stronger evidence - and recent work shows exactly that.
Models build the structure they need from context
It would be one thing if large models simply carried a library of fixed structures learned during training. The more striking finding is that they build or reconfigure structure from the current context.
In Context Is King, researchers gave models arbitrary, meaningless tokens and declared relationships between them in the prompt. Depending on the rule, the same tokens were organised internally into a cycle or a branching tree. When the in-context specification contradicted a strong pretrained prior, the context-set geometry won in the more capable models. And activation patching showed the model was actually using that geometry - swap one entity’s activation for another’s, and the answer changed accordingly. The authors deliberately left open whether the geometry is built fresh or a stored one is reconfigured. But operationally, it’s the one the context specifies.
Large Language Models Develop Belief State Geometry In-Context goes further. Its authors prompted six open models, from 2B to 9B parameters, with long sequences from hidden Markov models the models had never seen, and asked what they would need to represent in order to predict them. The activations encoded something very close to the Bayesian belief state (a posterior over the hidden process generating the sequence) with linear probes recovering it at R² values between 0.83 and 0.99. Patching or steering that subspace moved the model’s predictions towards the injected belief, while random controls did not.
These are controlled settings, chosen for clean structure, and the models are modest in size. But the picture they paint is hard to square with “System 1 only”. Over the course of a context, the model infers a hidden generative process, maintains a structured belief over it, and computes its predictions through that belief. That’s a small model being constructed and used inside a larger one. And notice that none of this changed a single weight - it all happens in context, which will matter later in our discussion.
So why are LLMs still so fragile?
If models build and use latent models, why can a reworded prompt send their logic off a cliff? Why can a model clearly demonstrate that it knows something, then seem to forget it three paragraphs later?
I think this is where a lot of the “LLMs don’t reason” evidence gets misread. The mistake is assuming that if a useful structure exists, it must always control behaviour. It doesn’t. As I explored in What Makes LLMs So Fragile (and Brilliant)?, a model is constantly balancing three kinds of internal influence. There’s a compact state written mid-stack and cheaply reused - like jotting a key fact on a notepad. There are routes that recompute what they need on the fly - flexible, but sensitive to order and wording. And there are anchors, early tokens such as system prompts that later layers keep looking back to - like keeping a finger on the first page of a book. All three meet in the residual stream and compete, token by token, to shape what happens next. I call that competition arbitration.
That gives us four distinct things that can each be true or false:
a latent model exists
the current context recruits it
it is integrated into the model’s working state
it wins enough of the arbitration to actually control behaviour
A failure at step four does not mean step one failed.
Sometimes the model has the right structure and still does the wrong thing.
The recent papers let us watch this competition directly. In the belief-state study, an injected belief that conflicted with the rest of the context struggled while only a few positions were intervened on, because the untouched context could still compete - intervene on more positions and the injected belief took over. Interventions in early layers were often simply repaired by the network, which reinstated the original prediction. Context Is King shows the same contest between a context-set structure and a pretrained prior, with the outcome depending on the model.
Anthropic’s work on what they call J-space then adds an integration step. In Verbalizable Representations Form a Global Workspace in Language Models, they identify a small set of representations that the model can report, deliberately summon and hold, use for silent intermediate reasoning, and pass to arbitrary downstream computations - functionally, this is a global workspace. Equally telling, most of what a model does (fluent text, simple recall, grammar) runs without it. So the picture isn’t that there’s one master model. Instead it’s weights and context recruiting local latent models, a workspace integrating selected pieces of them, and arbitration deciding which of them wins to drive the next step.
In reality, that is not comforting if you’re deploying these systems in medicine or some other critical domain. But “the right structure sometimes loses” is a very different diagnosis from “there is no reasoning machinery here at all” - and these different diagnoses lead to different engineering.
A note on chain of thought
Graepel also points out that chains of thought are often concocted after the fact - the model reaches an answer by one route and reports another. The evidence for that is real, and nothing in this argument depends on traces being faithful.
The distinction that matters is between faithful as a report and causal as state. A trace can misdescribe the computation that produced it and still condition the computation that follows, because in an autoregressive system every emitted token becomes input to the next step - which is exactly what agentic-coding harnesses exploit, at scale. And the evidence in the papers above don’t rest on what models say about themselves. It rests on internal geometry that was measured, and then intervened upon.
The model is not the machine
I think Graepel’s strongest point isn’t that LLMs can’t produce useful intermediate reasoning. It’s that they don’t maintain the epistemic machinery we’d want from a trustworthy scientific reasoner - what is believed, what is doubted, which hypotheses compete, what evidence bears on each, what to test next, and why a belief changed. That’s a very good engineering goal. But there’s a hidden assumption in asking the neural network itself to supply all of it.
We humans don’t work that way either. In 1998 Andy Clark and David Chalmers published The Extended Mind, built around Otto, who relies on a notebook as his memory. If the notebook is reliably available and tightly woven into how Otto acts, they ask, why treat what he retrieves from it differently from what someone else retrieves from biological memory? You don’t need to accept every philosophical implication to see the practical point of this argument. Scientists keep lab books. Programmers use tests, debuggers and version control. Teams spread cognition across people with different expertise. We externalise state constantly, and very often that’s precisely what makes our reasoning reliable.
Otto making notes in a notebook doesn’t mean “Otto doesn’t reason”. The previous sections present the case that the model, like Otto, does real reasoning of its own - and that is what makes it worth building around. In a real sense it already keeps a notebook on the inside, in those anchors held in its key-value cache. But the real question in the Extended Mind analogy is then what we add on the outside.
So I think we need to change the object we’re looking at. The modern large model placed inside a harness is not a finished product. It is a kernel for a general inference platform, wrapped in context, persistent memory, retrieval, code execution, search, tests, critics and subagents. The model interprets the situation, builds local latent models, proposes hypotheses, decides which tools might help, and integrates the results. It doesn’t have to own arithmetic, persistent records, version history, or the verdict on whether the code actually passed.
Where we diverge
Graepel sees this landscape clearly. He writes that recent advances in LLMs and other neural models give us the capability to take on general reasoning problems - they can suggest ways of tackling a problem, interact with tools through APIs or code, and help assess whether a claim is supported by evidence. Then comes the condition that matters most to him. To keep the system honest, “an independent part of the system” must evaluate each move by how much it actually resolves uncertainty, with beliefs updated only when evidence backs the change.
I agree with all of that. Where we differ is what follows from it. As I read him, he sees these capabilities as raw materials that still need substantially new machinery (perhaps a fundamentally different architecture) before they amount to reasoning. His closing line places today’s systems on the side of “a convincing story told after the fact”, rather than an auditable sequence of evidence, inference and belief revision.
My reading is that this machinery is already available and can be assembled, from exactly the parts he lists. These are two reasonable bets. His comes from having built one of the most important reasoning systems in the history of AI. Mine comes from building harnesses around these models every day. So rather than argue it in the abstract, here’s a concrete example.
Building the missing machinery
In agentic coding I rarely ask one model to propose a solution, critique itself, implement it and declare victory. I split the work. A developer agent states a hypothesis, predicts what should happen, writes a test and makes the change. A critic agent tries to break it - hunting for hidden assumptions, alternative explanations, weak tests and edge cases. Then the harness runs the tests. Neither model gets to decide whether the code passed. Reality gets a vote. Using different models or providers can strengthen this further - not as majority voting, since more agents agreeing doesn’t make anything true, but to reduce correlated failure and turn disagreement into better tests. None of this is a comment on what a single model can do, any more than peer review is a comment on what a single researcher can do. Separating proposal from challenge is simply how diverse perspectives make reasoning more robust.
To make this tangible I’ve formalised this pattern into an epistemic ledger skill, which I’ve published on GitHub. It’s based on Karpathy’s LLM Wiki concept, customised to this particular need. It keeps a persistent, Git-versioned Markdown record of five things:
evidence (what was actually observed)
hypotheses (candidate explanations)
claims (propositions other work is currently authorised to depend on)
questions (what remains materially unresolved) and
decisions (actions taken on the current state).
Observations are kept apart from interpretations, so a model’s inference can’t simply harden into a fact. Hypotheses state their predictions and what would falsify them. Critics are asked to falsify, never to verify. And evidence comes from executors (tests, tools, measurements) not from anyone’s account of their own reasoning, which is another reason the faithfulness of chain of thought stops mattering here.
Here’s how that maps onto what Graepel says a trustworthy reasoner needs:
One design choice is worth drawing out. Claims in the ledger are not certified truths. A claim is a revocable licence to depend on something - it carries an explicit scope, known limitations, and the conditions under which it should be revisited, and it is demoted the moment new evidence contests it. Graepel writes of a system that can “accumulate certified knowledge”, and I suspect he means something very similar. But the word invites a finality that science itself never grants, and the ledger deliberately avoids it.
He also wants a system that improves its reasoning policy by learning from past reasoning experiences. That doesn’t have to mean changing weights. As the research above shows, a model’s latent machinery adapts to what’s in its context. An optimised skill and a well-kept ledger both shape what gets recruited on every run, while the persistence lives in the skill and the ledger themselves. The policy that improves belongs to the extended system.
The skill also states a principle I think applies well beyond software: “The ledger models what is known, not the topology of the agents reasoning about it”. Swap models, add critics, change providers - the epistemic state stays put.
None of this makes the hard problems disappear. A critic can be wrong. A test can measure the wrong thing. Two models can share the same blind spot. And the ledger’s rules are enforced softly (by instructions, review and history) so an agent can still record an inference as an observation unless something stops it. But stronger enforcement is easy to add: schema checks, commit hooks that reject a claim without its dependencies, CI that validates the record graph, external audits. Literally anyone could prompt an agentic-coding harness to build those in minutes or hours. And the patterns here are simply the ones I’ve chosen in order to make this tangible - anyone can design their own to reach the same goals.
These are real engineering problems. But notice how different they are from the claim that we lack the machinery for reasoning because an LLM generates one token at a time. We already have the pieces. The real challenge is in how we choose to compose them.
AlphaGo was already an extended reasoning system
There’s a nice symmetry here, and it’s Graepel’s own. AlphaGo was never “just a neural network” - that’s the whole point of his opening. Its power came from learned intuition embedded in a larger system with explicit state and a mechanism for checking itself. Perhaps the deeper lesson isn’t that LLMs must become search algorithms internally, but that learned intuition becomes far more powerful when it’s embedded in a system like that.
Go handed AlphaGo a beautifully constrained world and a well-defined tree. Science, software engineering and open-ended research don’t. The state is messier, the actions are open-ended, evidence arrives from tools, experiments, documents and other people, and the questions themselves change as we learn. A fixed search algorithm struggles when nobody has defined the search space. An LLM can help construct it as it goes (proposing hypotheses, noticing what’s missing, writing the experiment, calling the tool, interpreting the result, and handing the conclusion to a critic) while the harness keeps the state that shouldn’t be trusted to transient activations alone.
That looks less like System 1 pretending to be System 2, and more like a new kind of modular reasoning platform - where the model is the kernel.
So, do LLMs reason?
I think the answer has two parts.
By ordinary definitions, the model itself does reason. We can find structured latent representations inside it, intervene on them and watch behaviour change, see it build relational geometry and belief states from context, and find a workspace where selected representations are integrated and made available to the wider computation. Its fragility looks like what you’d expect when useful latent models have to be recruited, integrated, and keep winning arbitration. It does not look like their total absence.
But Graepel sets a higher bar - reasoning in a way that a scientist might recognise, with persistent, inspectable, evidence-gated belief revision. That’s a legitimate bar for the work he cares about, even though it just folds reasoning into the more strict auditable reasoning. By that stricter definition, the extended system can reason too, and we can make large parts of it inspectable. But most interestingly of all this is possible precisely when we don’t force all of it to live inside the network.
That may be the real opportunity hiding in his critique. The machinery he says is missing does need to exist. I’m just not convinced we have to wait for a fundamentally different neural reasoner to get it. We can already build it around the models we have - and I’d genuinely welcome his view on where that really falls short.
And I hope you’ll download the SKILL.md and try it for yourself too.



