Summary
Rishi Desai, ML engineer at Abundant AI, presents SWE-Marathon at the AI Engineer conference — a new benchmark designed to evaluate coding agents on project-scale tasks that run for multiple hours and consume up to a billion tokens. While existing benchmarks like Human Eval and SWE-bench test individual functions or GitHub issues, SWE-Marathon asks agents to build a full Slack clone, rewrite an entire JAX codebase in PyTorch, or implement a C compiler in Rust — tasks representing hundreds of hours of human engineering work.
Verification at this scale is the central engineering challenge, and SWE-Marathon addresses it with layered independent checks: hidden unit tests, reference parity checks, anti-cheating tests, and a computer-use agent (CUA) that evaluates full-stack products by actually operating them through a browser UI. This prevents agents from gaming the reward signal by probing the verifier rather than completing the intended work.
Leaderboard results show Claude Opus 4.8 with Claude Code as the top configuration at just 26% resolution — solving roughly one in four tasks — with GPT-4.5 plus Codex at 12% and substantially lower cost. Average trial length was 31 million tokens; the longest rollout consumed 877 million tokens. Desai concludes that agent scaffolding design matters as much as raw model capability, and that end-to-end autonomous project ownership remains far from solved.
📺 Source: AI Engineer · Published July 07, 2026
🏷️ Format: Deep Dive







