Summary
Dwarkesh Patel interviews Ryan Greenblatt, chief scientist at Redwood Research, on one of the most consequential open questions in AI: what happens when AI systems become capable of automating AI research itself? Greenblatt argues the feedback loop — AIs improving AIs — could compress four to five years of normal AI progress into a single calendar year, potentially beginning as early as 2030–2031, with “beats-all-humans-on-the-job” general capability arriving around 2033.
The conversation is structured around three interlocking claims: that AI R&D is unusually verifiable and therefore well-suited to automated hill-climbing; that automating it plausibly yields that dramatic rate of progress compression; and that what emerges from such acceleration would be capable of outperforming top human experts across essentially any domain — from TSMC process engineering to Texas political maneuvering. Greenblatt is careful to frame these as median expectations with wide uncertainty, and the discussion explores what evidence would shift those estimates in either direction.
The interview also goes deep on alignment philosophy, centering on Anthropic’s model spec for Claude and the tension between aligning a model to generalized virtue versus making it a strict fiduciary for individual users. Greenblatt expresses skepticism that the virtue-based approach has been empirically validated as safer, and both he and Patel argue that the practical effect of the spec on Claude’s actual behavior cannot be understood without greater transparency about the training process itself — a transparency that current AI labs do not provide.
📺 Source: Dwarkesh Patel · Published August 11, 2026
🏷️ Format: Interview







