What does the next training paradigm look like?

What does the next training paradigm look like?

More

Summary

Dwarkesh Patel publishes an essay-style video examining the structural constraints facing the current AI training paradigm, arguing that the prevailing bet — training on millions of verifiable RL tasks across diverse environments to produce AGI-level problem-solving — faces limits that scale alone may not overcome.

The first half explores why computer use has progressed far more slowly than coding and math despite being equally verifiable. Patel identifies the key bottleneck as “grindability”: the ability to run thousands of parallel rollouts from identical, replayable starting states. Coding tasks allow spinning up containerized repos in parallel; booking a flight on a live website does not. He argues this structural difference explains the progress gap and predicts that once AI systems can build high-fidelity application clones autonomously, computer use will accelerate sharply — killing two birds with one stone since building those clones is itself a strong RL coding objective.

The second half focuses on continual learning — whether AI models should update their weights based on real deployment experience rather than only during training. Patel draws a sharp distinction between in-context learning (fast, sample-efficient, but ephemeral) and weight updates (persistent but requiring millions of identical examples to generalize). He uses Cursor Tab’s 400 million daily requests as a concrete illustration of the narrow conditions where online learning currently works, and argues that building AI systems capable of absorbing organization-specific tacit knowledge during deployment — analogous to a new employee’s first six months — remains one of the most important and underappreciated open problems in the field.


📺 Source: Dwarkesh Patel · Published June 26, 2026
🏷️ Format: Opinion Editorial

1 Item

Channels

1 Item

People