OpenAI researcher on agent swarms & recursive self-improvement

OpenAI researcher on agent swarms & recursive self-improvement

More

Summary

Dwarkesh Patel interviews Noam Brown, an OpenAI researcher who helped develop the o1 reasoning models and now works on multi-agent systems, following OpenAI’s announcement that a swarm of 10,000 AI agents solved one of the Millennium Prize Problems using 130 billion tokens over 88 hours. Brown unpacks how reasoning models trade inference-time compute for better performance, and why multi-agent parallelization — used in OpenAI’s “Ultra Mode” with configurable agent counts — offers a way around single-agent latency bottlenecks.

The conversation includes concrete scaling data: running four agents in parallel roughly halves completion time at double the compute cost, with diminishing but still meaningful returns up to 16 agents, and effects varying significantly by task type, with math and web search proving especially parallelizable. Brown and Patel debate whether this points toward a gradual intelligence explosion or a slower, jagged path to more general capability, and discuss how narrow improvements in learning efficiency could still generalize broadly.

The discussion turns toward longer-term implications, including projections that by the end of the decade leading AI labs could have enough compute to run the equivalent of hundreds of millions of human-level intelligences. It’s a technically grounded look at where multi-agent reasoning systems stand today and where they may be heading, straight from someone building them at OpenAI.


📺 Source: Dwarkesh Patel · Published September 17, 2026
🏷️ Format: Podcast

1 Item

Channels

1 Item

Companies