AI researchers debate how close we are to recursive self-improvement

AI researchers debate how close we are to recursive self-improvement

More

Summary

Dwarkesh Patel hosts a roundtable with three researchers from frontier-adjacent AI labs: Beren Millidge, CTO of Zyphra (an open-source model developer); John Schulman, chief scientist at Thinking Machines and previously a co-founder of OpenAI who led the RLHF work behind ChatGPT; and Charlie O’Neill, head of model training at Baseten. The central question is what technical factors could prevent a superintelligent AI transition by 2036 — and how close recursive self-improvement actually is.

The conversation covers several serious technical arguments: whether current transformer-plus-RL architectures are near a global optimum or require a paradigm-level discontinuity to continue improving; why even a model slightly better than all humans at AI research could trigger explosive self-improvement once run at scale across millions of parallel instances; and whether the 50% of inference compute currently spent on deployment (rather than training) represents a latent intelligence explosion waiting to happen once models can meaningfully learn from deployment data.

Schulman draws on his direct experience building RLHF to discuss the cycle of models feeling transformative on release and then mundane after a month of use, and why that cycle has so far prevented explosive capability growth despite rapid progress. Millidge raises the analogy of Moore’s Law requiring repeated discrete innovations to maintain the appearance of a smooth scaling curve. The discussion is notably candid and technically grounded, making it a strong reference for anyone tracking the state of expert opinion on AGI timelines and self-improvement dynamics.


📺 Source: Dwarkesh Patel · Published September 11, 2026
🏷️ Format: Podcast

1 Item

Channels