Summary
Dwarkesh Patel presents a detailed analytical argument about what is actually driving AI progress — and his answer is data, not architecture or training tricks. The central claim is that sample efficiency (how much data a model needs per skill) has not meaningfully improved over recent years; what has changed is that labs are collecting vastly more and better training data, particularly through reinforcement learning pipelines that function as large-scale synthetic data generators. Patel frames RL with verifiable rewards (including GRPO) as fundamentally a method for identifying which data is worth keeping, then training the model on correct rollouts.
He quantifies the gap with specific comparisons: humans absorb roughly 200 million language tokens by adulthood, while frontier models train on tens to hundreds of trillions — approaching a millionfold difference. Every marginal skill requires enormous volumes of domain-specific human expert demonstrations, and Patel cites Mercor and Surge job listings for Word document specialists, M&A lawyers, and market research consultants as evidence of how labor-intensive and domain-specific that data pipeline is.
A structural argument with practical implications: open-source models lag frontier models by roughly four months (per Epoch AI research), and Patel argues this is because training data can be distilled from public APIs while proprietary hyperparameters and training tricks cannot. His conclusion is that the invisible data black hole — not model architecture — is the real competitive moat in AI, and it is becoming increasingly expensive to fill.
📺 Source: Dwarkesh Patel · Published June 19, 2026
🏷️ Format: Opinion Editorial







