Summary
Pierluca D’Oro, founder of Programmabs and former researcher at Meta Superintelligent Labs, delivers a conference talk exposing fundamental statistical flaws in how computer use agents are evaluated today. The central insight: a “replay agent” — a simple script that records successful action sequences from a frontier model and blindly replays them — achieves the same or better benchmark scores as the frontier models it was derived from on standard benchmarks like OSWorld and MobileWorld. This is not a quirk but a consequence of benchmark determinism, and D’Oro proves formally that the widely used pass@K metric is mathematically equivalent to evaluating a replay agent under these conditions.
To address these problems, D’Oro presents the PRISM principles — a design framework for building trustworthy agent evaluation environments: environments should be multi-factorial (stochastic across parameters like data, appearance, and initial state), should include verifiers to confirm that generated combinations are valid, should be sandboxed, and should faithfully reproduce real-world systems. He describes DIG, a benchmark built to these principles that uses a compiler-like system to generate verified task configurations, preventing replay-agent exploitation while enabling billions of valid combinations.
The talk demonstrates concrete results showing that frontier models fail to maintain performance robustly across variation axes, with worst-case performance dropping significantly. For anyone building or evaluating computer use agents, the work is essential context for why leaderboard scores on existing benchmarks may be substantially misleading — and what a more rigorous evaluation regime should look like.
📺 Source: AI Engineer · Published August 14, 2026
🏷️ Format: Deep Dive







