Stop Evaluating Models Like It’s the 50s – Alejandro Vidal, Mindmakers

Stop Evaluating Models Like It’s the 50s – Alejandro Vidal, Mindmakers

More

Descriptions:

Alejandro Vidal, founder of Mind Makers, argues that the AI industry’s standard approach to LLM benchmarking—summing correct answers and treating every question as equally weighted—is a relic of 1950s classical test theory. Drawing on his dual background in psychology and computer science, he introduces Item Response Theory (IRT) as a more rigorous alternative for measuring model intelligence.

Using real benchmark data from epoch.ai, Vidal demonstrates how IRT assigns each question a difficulty parameter (B) and a discrimination parameter (A), while estimating each model’s latent intelligence score (theta). This framework surfaces hidden pathologies in existing benchmarks: mislabeled gold answers, questions negatively correlated with actual capability, and redundant items that add evaluation cost without adding signal. He walks through a live example where a benchmark’s canonical answer conflates total passengers with total casualties—a subtle labeling error that skews model rankings.

The practical payoff is substantial. By selecting only high-discrimination items, Vidal achieves 99% correlation with the original model ranking using just 97 questions instead of 484—a nearly 5x reduction in benchmark size and token cost. Random item selection performs far worse at the same compression ratio, validating the IRT-guided approach. For teams running private benchmarks to choose between open-source models, this methodology offers a principled way to cut evaluation cost while maintaining ranking fidelity.


📺 Source: AI Engineer · Published July 13, 2026
🏷️ Format: Deep Dive

1 Item

Channels