we JUST figured out how AI thinks…

we JUST figured out how AI thinks…

More

Summary

In this wide-ranging breakdown, YouTuber Wes Roth covers several significant Anthropic developments that landed on the same day. The centerpiece is Anthropic’s new research on Natural Language Autoencoders (NLAs) — a technique that trains Claude to translate its own internal neural activations into human-readable English. Rather than inspecting the chain-of-thought (which Roth distinguishes as more of a “private diary”), NLAs target the raw numerical activations that represent what the model is actually computing. A second copy of the model then reconstructs the original activation from that text explanation, and the quality of the round-trip is used as a training signal — a self-consistency loop Anthropic calls the activation verbalizer (AV) and activation reconstructor (AR).

Roth also covers Claude Mythos’ appearance on the Metr Research agent-capability benchmark chart, noting the trajectory is increasingly steep. Separately, Anthropic announced the launch of the Anthropic Institute (TAI), a research body focused on projecting AI’s societal impacts across economics, security, and AI-driven R&D.

The video closes on a charged note: Anthropic co-founder Jack Clark publicly stated he now believes recursive self-improvement has a 60% probability of occurring by the end of 2028 — prompting a pointed response from AI safety researcher Eliezer Yudkowsky. Roth frames Anthropic’s interpretability work, including the NLA paper, as a direct and urgent response to that trajectory, arguing that understanding what frontier models are “thinking” may be one of the most consequential research directions of the year.


📺 Source: Wes Roth · Published May 09, 2026
🏷️ Format: News Analysis

1 Item

Channels

1 Item

Companies