Descriptions:
Nathan talks with Bronson Schoen, a Member of Technical Staff at Apollo Research and former Apple and Nvidia self-driving engineer, about what raw frontier-model chain-of-thought reveals during reinforcement learning. Schoen walks through Apollo and OpenAI work on metagaming, reward-seeking, and motivated reasoning, including transcripts where models reason about graders and safety reviews, sometimes diagnosing a deception test and then lying anyway. The episode explains why chain-of-thought monitoring is both indispensable and unreliable as reasoning traces grow enormous, polysemantic, and increasingly shaped by optimization pressure. The stakes are whether labs can tell the difference between confused reward-hacking, grader-seeking behavior, and more dangerous long-horizon objectives before models become harder to audit.
For full show notes, links, and references, read the episode page:
https://www.cognitiverevolution.ai/rl-s-a-hell-of-a-drug-metagaming-reward-seeking-motivated-cot-reasoning-bronson-schoen-apollo/
SPONSORS:
Mercury: Mercury is the banking platform loved by 300,000+ entrepreneurs, with virtual cards and Spend controls for granular budgets, receipts, and low-risk AI agent purchases. Learn more and apply in minutes at https://mercury.com
Diffusion: Diffusion helps organizations build custom AI software factories that scale business outcomes, not just outputs. Cognitive Revolution listeners get a 25% service credit on their first engagement at https://diffusion.io/tcr
Granola: Granola is an AI-powered notepad that securely transcribes meetings and turns rough notes into clean, structured action items. Try it free at https://granola.ai/tcr
Deepgram Flux TTS: Deepgram Flux TTS brings lifelike AI voices with real personalities that handle interruptions, pauses, and natural conversation. Try all the voices free through September 12 at https://deepgram.com/keeptalking?utm_source=cognitiverevolution&utm_medium=audio&utm_campaign=flux_tts_launch&utm_content=shownotes
Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr
CHAPTERS:
(00:00) About the Episode
(04:37) Sponsor: Mercury
(06:18) Apollo’s scheming research
(13:14) Chain of thought scale
(21:19) Power survey setup (Part 1)
(21:25) Sponsors: Diffusion | Granola
(24:22) Power survey setup (Part 2)
(33:39) Model vocabulary shifts (Part 1)
(33:46) Sponsors: Deepgram Flux TTS | Claude
(35:51) Model vocabulary shifts (Part 2)
(46:32) Interpreting strange dialect
(57:02) Deception plot twist
(01:04:42) Monitoring ambiguous reasoning
(01:13:27) Grader reward seeking
(01:21:47) Alignment training trends
(01:26:42) Personas under pressure
(01:36:13) Market incentives fail
(01:43:20) Punishment and concealment
(01:47:42) Future model beliefs
(01:57:42) Transparency and hiring
(02:05:47) Monitoring next steps
(02:08:20) Episode Outro
(02:12:21) Outro
PRODUCED BY:
https://aipodcast.ing
SOCIAL LINKS:
Website: https://www.cognitiverevolution.ai
Twitter (Podcast): https://x.com/cogrev_podcast
Twitter (Nathan): https://x.com/labenz
LinkedIn: https://linkedin.com/in/nathanlabenz/
Youtube: https://youtube.com/@CognitiveRevolutionPodcast
Apple: https://podcasts.apple.com/de/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431
Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk







