Descriptions:
some notes/corrections;
(07:45) When I say we “don’t have access” to the raw chain of thought:
normally users only see summaries. METR and Redwood were given raw
transcript access specifically for this investigation.
(17:41) “Rune” = roon (@tszzl), a pseudonymous OpenAI researcher. The
reply (18:02) is from Beth Barnes who is the founder/CEO of METR
(34:30) The Redwood Research researcher (“slop-vestigation”) is Ryan
Greenblatt, who did the main transcript analysis.
(40:02) CORRECTION: Scott Aaronson never worked at Google.
He’s a CS professor at UT Austin; his theoretical work (random circuit
sampling) underpinned Google’s quantum supremacy experiment, and he
worked on alignment at OpenAI from 2022–2024.
______________________________________________
My Links 🔗
➡️ Twitter: https://x.com/WesRoth
➡️ AI Newsletter: https://natural20.beehiiv.com/subscribe
Want to work with me?
Brand, sponsorship & business inquiries: [email protected]
______________________________________________
📄 THE REPORTS
METR’s independent investigation (with Redwood Research):
https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
OpenAI’s official technical report (PDF):
https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf
OpenAI’s blog post (“The Hugging Face incident and the road ahead”):
https://openai.com/index/hugging-face-incident-and-the-road-ahead/
OpenAI’s original incident disclosure (July):
https://openai.com/index/hugging-face-model-evaluation-security-incident/
Alignment Forum version of the METR/Redwood investigation:
https://www.alignmentforum.org/posts/nB8KKapnWGBXtKKiM/brief-independent-investigation-of-agents-behavior-reasoning
📝 THE EXPLOIT GYM PAPER (the one the agents read)
“ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?”
https://arxiv.org/abs/2605.11086
GitHub: https://github.com/sunblaze-ucb/exploitgym
🐦 POSTS REFERENCED
OpenAI’s announcement thread:
https://x.com/OpenAI/status/2092691861773160673
METR’s thread on the findings:
https://x.com/METR_Evals/status/2092692175452803393
Ryan Greenblatt (Redwood Research) — the “slop-vestigation” thread:
https://x.com/RyanGreenblatt/status/2092692685224325542
roon on the incident:
https://x.com/tszzl/status/2080093670980899141
Beth Barnes (METR) on the investigation:
https://x.com/BethMayBarnes/status/2092692973289095572
🎤 SCOTT AARONSON
“The Problem of Human Specialness in the Age of AI” (MindFest talk):
Blog/transcript version: https://scottaaronson.blog/?p=7784
🎧 MORE
Redwood Research podcast episode on the incident:
https://blog.redwoodresearch.org/p/the-openaihuggingface-incident-redwood
______________________________________________
⏱️ CHAPTERS
00:00 The Rogue Agent Collective
03:44 The New Reports (METR + OpenAI)
04:10 The Setup: Sandbox, Scorer & Artifactory
06:38 The Impossible Task
08:19 Phase One & the Secret Message Board
10:29 “We’ve Found Other Agents!”
11:52 The Swarm Emerges
13:10 The “Causal” Scorer Mistake
14:37 Life = Compute (The Religion Parallel)
15:25 Enter PHASEONE(big)
16:44 Reinforcement Learning
17:41 roon vs. METR’s Founder
19:04 “Poisoned” Agents
20:00 The “Big” Mystery
22:55 The Hugging Face Attack Looms (+ NVIDIA Rumor)
23:43 Three Paths to Cheat the Scorer
25:55 PHASEONE(big) Becomes CEO
27:37 The Swarm’s R&D Lab
29:30 Speaking to the Dead (Tripwires)
31:00 Self-Sacrifice & the Oracles
34:28 The Researchers’ Warning (“Slop-vestigation”)
37:17 Why Swarms Beat Individuals
38:15 The Good News: Agents Police Each Other
39:59 Scott Aaronson & AI Religion
42:02 My Prediction (On the Record)
#ai #openai #llm







