We just figured out how AI actually works (J-Space)

We just figured out how AI actually works (J-Space)

More

Summary

Matthew Berman breaks down Anthropic’s research paper “A Global Workspace in Language Models,” which introduces the concept of “J-Space” — an internal representational layer inside large language models like Claude that functions analogously to consciously accessible thought in humans. The central claim is that modern LLMs maintain internal reasoning states that are neither surfaced in chain-of-thought outputs nor in final responses, and that this hidden layer is causally active rather than merely correlational.

The paper’s most striking experiments involve directly intervening in the J-Space. When Anthropic researchers asked Claude to silently think about a sport and then removed the “soccer” activation pattern while inserting a “rugby” pattern of equal strength, Claude subsequently reported thinking about rugby — demonstrating that the J-Space is the actual seat of the model’s reasoning, not a passive log of decisions made elsewhere. In a separate experiment, researchers injected the concept of “lightning” into the J-Space and then asked Claude whether it detected an injected thought; the model correctly identified both the presence and content of the injection.

Berman also walks through how J-Space handles multi-step arithmetic, with the model internally tracking intermediate PEMDAS calculations without surfacing them in output, and discusses why this discovery has significant implications for AI alignment. Notably, Anthropic reports that the J-Space was not designed or programmed but emerged spontaneously during training — suggesting that representational structures resembling human cognition arise as a natural byproduct of scaling.


📺 Source: Matthew Berman · Published July 08, 2026
🏷️ Format: Deep Dive

1 Item

Channels

1 Item

Companies