Inside the Race to Measure Frontier Intelligence

Inside the Race to Measure Frontier Intelligence

More

Summary

This a16z interview explores the growing market for independent AI evaluation, featuring the founder of Vows, a benchmarking company launched in 2024 after the founding team concluded that public benchmarks were no longer sufficient to measure frontier model progress. The conversation opens with a striking example: when Meta released Llama 4, it posted impressive scores on open public benchmarks but significantly underperformed on Vows’ private held-out evaluations — a gap the guest attributes to labs optimizing directly on benchmarks whose questions and rubrics are publicly available.

The discussion covers why labs cannot credibly self-report capability gains, drawing analogies to rating agencies and audit firms in other trillion-dollar industries. The guest argues that independent evaluation is becoming existential for enterprises as well as labs — illustrated by a Fortune 10 company that initially capped Claude Code usage at $100 per engineer per day (later raised to $300), creating productivity dead zones when rate limits hit each afternoon. The conversation also touches on how evaluation complexity is rising as workflows grow more sophisticated, why model routing companies like OpenRouter (recently acquired by Stripe) depend on high-quality evals to function, and where government safety evaluation fits alongside commercial benchmarking.


📺 Source: a16z · Published September 09, 2026
🏷️ Format: Interview

1 Item

Channels