State of Data — Sean Cai, Independent / State of Data

State of Data — Sean Cai, Independent / State of Data

More

Summary

In this AI Engineer conference talk, independent analyst Sean Cai delivers a dense market analysis of the AI training data industry — what he argues is the most underfunded and least understood input in the foundation model supply chain. Cai reframes data not as static annotation work (the $10–15B labeling market most people picture) but as a dynamic supply chain that has fragmented into 20–30 specialized vendors per lab, with quality no longer scaling linearly with quantity.

A central concept is the distinction between “type one” data (authentic process traces like GitHub commits or session replays) and “type two” data (contrived expert-generated tasks), and why the type-two playbook that worked for GPQA-style benchmarks in 2024 breaks down on long-horizon, unverifiable work. Cai is particularly sharp on benchmark integrity: he describes a structural conflict of interest in which vendors generate plausible tasks, cherry-pick divergences, package them as hard benchmarks, then sell data to hill-climb those same benchmarks — what he calls Goodhart’s Law with a profit motive. Cross-harness and cross-infrastructure differencing, he argues, is the primary driver of benchmark divergence across frontier evaluations.

The talk is aimed at practitioners, researchers, and investors trying to understand where model improvement actually comes from and why the data layer represents both a significant market opportunity and a source of widespread misinformation in AI capability claims.


📺 Source: AI Engineer · Published July 26, 2026
🏷️ Format: Keynote Launch

1 Item

Channels

2 Items

Companies