Summary
At the AI Engineer conference, Uri Rolls of Arithmetic and Thomas Wolf of Hugging Face unveil MASOV — a new benchmark designed to test whether AI models can reason through real-world cybersecurity access-control vulnerabilities. The talk positions the benchmark alongside ARC-AGI 3 in terms of the dynamic world-modeling demands it places on models, arguing that cybersecurity represents a far wider exploration field for AI than is commonly recognized.
The presenters explain that current frontier models achieve only 1–2% success rates on MASOV’s tasks, which involve reasoning across large, chained application environments without access to source code or the internet. The benchmark focuses on access-control flaws — ranked number one on the OWASP list and underpinning a $30 billion security industry — because they are logic-based rather than syntactic, requiring models to build and reason from an internal world model on the fly rather than pattern-matching known vulnerability signatures. Every step in the exploitation and discovery chain is deterministically graded.
A central theme is the role of open-source models in cybersecurity defense. Wolf and Rolls challenge the assumption that closed models are inherently better for security, arguing that open-weight models will be essential to scalable defensive AI. Their evaluation environments are built from original zero-day exploits found by the team’s own vulnerability researchers in widely distributed open-source software — a methodology designed to prevent data contamination and force genuine out-of-distribution reasoning.
📺 Source: AI Engineer · Published July 24, 2026
🏷️ Format: Keynote Launch







