Stop Evaluating Models Like It’s the 50s – Alejandro Vidal, Mindmakers

Stop Evaluating Models Like It’s the 50s – Alejandro Vidal, Mindmakers

More

Summary

Alejandro Vidal of Mindmakers presents a conference talk arguing that LLM benchmarks are fundamentally flawed because they treat every question as equally weighted — an assumption modern psychometrics discarded decades ago. Drawing on Item Response Theory (IRT), he introduces a framework that estimates per-item difficulty (parameter B) and per-model latent ability (theta), producing richer comparisons than raw accuracy scores.

The framework surfaces surprising results: Claude Opus 4.1 and Gemini 3 Pro can share identical benchmark scores while differing by a full standard deviation in estimated ability, depending on which specific questions each model answered correctly. IRT-based residual analysis can also detect potential benchmark data leakage — flagging cases where a model answers an unexpectedly hard question correctly — and identify unstable model behavior, with o4 mini showing the most internal inconsistencies in the examples shown.

Vidal also demonstrates benchmark compression: by ordering items from most to least informative, a drastically smaller subset of questions can achieve 99% correlation with the full test’s model rankings. All benchmarks, tools, and datasets referenced in the talk are being shared publicly, and Vidal advocates for these psychometric techniques to become standard practice for both benchmark creators and practitioners selecting models for production use.


📺 Source: AI Engineer · Published July 12, 2026
🏷️ Format: Deep Dive

1 Item

Channels