Claude’s Brain Has A Secret… And Scientists Found It

Claude’s Brain Has A Secret… And Scientists Found It

More

Summary

Two Minute Papers host Dr. Károly Zsolnai-Fehér covers a research paper revealing how Claude and similar large language models spontaneously develop structured internal representations for positional reasoning—specifically for tasks like estimating whether a word will overflow a line—despite never being explicitly trained to do so and receiving no character count information in their inputs.

The key finding is that the model develops neuron-like features analogous to biological “place cells” discovered in mouse brains: individual neurons that fire only when the animal occupies a specific spatial location. In the AI’s case, the model constructs features that fire depending on how far along a line the current token falls, encoding position as a low-dimensional curved manifold—specifically a rippling spiral—so that each position gets its own distinct channel with minimal interference from neighboring values. Crucially, the model doesn’t count raw characters; it counts tokens and multiplies by roughly four, a heuristic it invented without being told.

The paper’s broader implication, which the video emphasizes, is that AI systems appear to build hidden internal tools for novel problems during training—tools no one designed and that only interpretability research can surface. Zsolnai-Fehér frames this as an early step toward what Isaac Asimov called “robopsychology”: the scientific study of machine minds from the outside in.


📺 Source: Two Minute Papers · Published July 15, 2026
🏷️ Format: Deep Dive

1 Item

Channels

1 Item

Companies