Understanding the inner thoughts of AI

Understanding the inner thoughts of AI

More

Summary

Google DeepMind’s official podcast brings together Professor Hannah Fry and Neil Nander, who leads the language model interpretability team at DeepMind, for a wide-ranging conversation on one of AI safety’s most active research fronts. The episode covers what interpretability research actually is, why it matters for making AI systems safe, and what the field has learned — and run up against — since its early optimistic results.

Nander introduces the core challenge using an evolutionary analogy: neural networks like Gemini are “grown, not designed,” emerging from millions of small gradient updates rather than explicit human specification. This makes reverse-engineering their internal logic a task analogous to what biologists do with evolved organisms. The conversation covers mechanistic interpretability specifically — the subfield that tries to identify meaningful structures inside model activations — tracing its history from Chris Olah’s early neuron-level discoveries at OpenAI through to modern probing techniques. Nander describes how the team found that emotional states like happiness and sadness correspond to geometric directions in activation space, detectable by training simple linear classifiers on model internals.

The episode also addresses one of the field’s most consequential open problems: building deception detectors for AI models. Nander explains why this is substantially harder than detecting sentiment — deception requires modeling the AI’s internal belief state, not just its output — and references a position paper his team published outlining the difficulties. For researchers, engineers, and anyone thinking about AI safety, this episode offers unusually direct access to the thinking of a senior researcher at one of the world’s leading AI labs.


📺 Source: Google DeepMind · Published July 10, 2026
🏷️ Format: Podcast

1 Item

Channels

2 Items

Companies