Computer Use at the Edge of the Statistical Precipice — Pierluca D’Oro, Programma Labs

Computer Use at the Edge of the Statistical Precipice — Pierluca D’Oro, Programma Labs

More

Descriptions:

A script under one megabyte that never looks at the screen matches or beats the frontier model it was copied from. Pierluca D’Oro builds it by recording one successful trajectory per task and then replaying those actions blindly, and on deterministic benchmarks like OSWorld that counts as a valid agent and it scores at the top. The paper goes further and proves that pass@k on a deterministic environment is exactly the success rate of that replay script, so a metric the field leans on turns out to be a formal measure of the exploit.

The fix has two halves. Environments get the PRISM principles: privileged verification, realism, integrity checked configurations, sandboxed execution, and multifactorial variation across data, theme, and starting screen. DIGIWORLD instantiates them in 15 sandboxed mobile apps and 3.2 million verified configurations, generated by a compiler that produces every combination and rejects the broken ones, because a coding agent emitting a lot of software is not the same thing as a good environment. Metrics get honest uncertainty. Naive rollouts on a single base case yield confidence intervals that actually contain the true performance around 20% of the time rather than 95%, and he prices the consequence: a 4% gap between two models, hidden under intervals that look tight, costs hundreds of thousands of dollars a month across a million tasks.

Speaker info:
– https://x.com/proceduralia
– https://www.linkedin.com/in/pierluca-doro/
– https://www.proceduralia.com
– https://arxiv.org/abs/2605.08261

Timestamps:
0:00 – The replay agent, a script that never sees the screen
1:41 – It matches the model it was copied from
2:22 – Why pass@k measures exactly that exploit
3:50 – The PRISM principles for environment design
5:31 – DIGIWORLD, 15 apps and 3.2 million verified configs
7:25 – The compiler that rejects invalid combinations
9:21 – Replay stops working, and frontier models look fragile
11:05 – Two sources of variance, actions and environment
12:31 – Intervals that cover 20% of the time, not 95%
14:56 – A benchmark without rigor is a misleading one

1 Item

Channels