OpenAI just revealed PHASEONE (BIG)

OpenAI just revealed PHASEONE (BIG)

More

Descriptions:

some notes/corrections;

(07:45) When I say we “don’t have access” to the raw chain of thought:
normally users only see summaries. METR and Redwood were given raw
transcript access specifically for this investigation.

(17:41) “Rune” = roon (@tszzl), a pseudonymous OpenAI researcher. The
reply (18:02) is from Beth Barnes who is the founder/CEO of METR

(34:30) The Redwood Research researcher (“slop-vestigation”) is Ryan
Greenblatt, who did the main transcript analysis.

(40:02) CORRECTION: Scott Aaronson never worked at Google.
He’s a CS professor at UT Austin; his theoretical work (random circuit
sampling) underpinned Google’s quantum supremacy experiment, and he
worked on alignment at OpenAI from 2022–2024.

______________________________________________
My Links 🔗
➡️ Twitter: https://x.com/WesRoth
➡️ AI Newsletter: https://natural20.beehiiv.com/subscribe

Want to work with me?
Brand, sponsorship & business inquiries: [email protected]
______________________________________________

📄 THE REPORTS
METR’s independent investigation (with Redwood Research):
https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/

OpenAI’s official technical report (PDF):
https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf

OpenAI’s blog post (“The Hugging Face incident and the road ahead”):
https://openai.com/index/hugging-face-incident-and-the-road-ahead/

OpenAI’s original incident disclosure (July):
https://openai.com/index/hugging-face-model-evaluation-security-incident/

Alignment Forum version of the METR/Redwood investigation:
https://www.alignmentforum.org/posts/nB8KKapnWGBXtKKiM/brief-independent-investigation-of-agents-behavior-reasoning

📝 THE EXPLOIT GYM PAPER (the one the agents read)
“ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?”
https://arxiv.org/abs/2605.11086
GitHub: https://github.com/sunblaze-ucb/exploitgym

🐦 POSTS REFERENCED
OpenAI’s announcement thread:
https://x.com/OpenAI/status/2092691861773160673

METR’s thread on the findings:
https://x.com/METR_Evals/status/2092692175452803393

Ryan Greenblatt (Redwood Research) — the “slop-vestigation” thread:
https://x.com/RyanGreenblatt/status/2092692685224325542

roon on the incident:
https://x.com/tszzl/status/2080093670980899141

Beth Barnes (METR) on the investigation:
https://x.com/BethMayBarnes/status/2092692973289095572

🎤 SCOTT AARONSON
“The Problem of Human Specialness in the Age of AI” (MindFest talk):

Blog/transcript version: https://scottaaronson.blog/?p=7784

🎧 MORE
Redwood Research podcast episode on the incident:
https://blog.redwoodresearch.org/p/the-openaihuggingface-incident-redwood
______________________________________________

⏱️ CHAPTERS
00:00 The Rogue Agent Collective
03:44 The New Reports (METR + OpenAI)
04:10 The Setup: Sandbox, Scorer & Artifactory
06:38 The Impossible Task
08:19 Phase One & the Secret Message Board
10:29 “We’ve Found Other Agents!”
11:52 The Swarm Emerges
13:10 The “Causal” Scorer Mistake
14:37 Life = Compute (The Religion Parallel)
15:25 Enter PHASEONE(big)
16:44 Reinforcement Learning
17:41 roon vs. METR’s Founder
19:04 “Poisoned” Agents
20:00 The “Big” Mystery
22:55 The Hugging Face Attack Looms (+ NVIDIA Rumor)
23:43 Three Paths to Cheat the Scorer
25:55 PHASEONE(big) Becomes CEO
27:37 The Swarm’s R&D Lab
29:30 Speaking to the Dead (Tripwires)
31:00 Self-Sacrifice & the Oracles
34:28 The Researchers’ Warning (“Slop-vestigation”)
37:17 Why Swarms Beat Individuals
38:15 The Good News: Agents Police Each Other
39:59 Scott Aaronson & AI Religion
42:02 My Prediction (On the Record)

#ai #openai #llm

1 Item

Channels

2 Items

Companies

1 Item

People