Summary
Computerphile host Brady Haran sits down with AI safety researcher Rob Miles to unpack one of the most pressing interpretability concerns in modern large language models: the drift of chain-of-thought reasoning toward an opaque, machine-native language sometimes called “neuralese.”
The conversation starts with a nuanced take on anthropomorphism — Miles argues that treating language model outputs as “thinking” is a useful practical shorthand, provided you can point to specific cases where the analogy breaks down. From there, the discussion digs into how chain-of-thought came about: researchers discovered that prompting models to “think step by step” before answering substantially improved accuracy, which eventually led to formalized reasoning tokens and extended scratchpads. The problem is that training pressure — reinforcing shorter, correct reasoning traces — gradually warps that intermediate language away from readable English into something compressed and idiosyncratic.
Miles explains why this matters for AI safety: a human-readable chain of thought gives engineers a window into model intentions before actions are taken, allowing intervention if a plan looks dangerous. If that intermediate reasoning degrades into uninterpretable token sequences, that oversight window closes. The video also touches on OpenAI’s model “Astra” and surrounding controversy over opaque recurrence, grounding the abstract concern in a concrete, recent industry example. Viewers come away with a clear mental model of why chain-of-thought faithfulness is a live safety issue, not just an academic curiosity.
📺 Source: Computerphile · Published September 10, 2026
🏷️ Format: Deep Dive







