Descriptions:
James Shi, founding engineer at Datacurve, presents DeepSWE (Deep Suite) — a contamination-resistant long-horizon software engineering benchmark designed to more accurately differentiate frontier AI coding models than existing alternatives like SWE-bench Pro.
Unlike benchmarks that mine tasks from closed public pull requests, DeepSWE’s 113 tasks are authored from scratch by open-source engineers who are active contributors to the repositories under test. The benchmark spans TypeScript, JavaScript, Python, Rust, and Go across nearly 100 distinct repositories (median one task per repo), closing the main contamination vectors that plague SWE-bench Pro — including the ability for models like Claude to run git log, identify golden patch commit hashes, and cherry-pick solutions. DeepSWE has replaced SWE-bench Pro in the Artificial Analysis coding agent index and has been cited by multiple frontier model labs. As of July 1st, 2026, Fable 5 holds the top leaderboard spot.
The talk also surfaces qualitative behavioral findings about how different models approach tasks: Claude tends to be exhaustive but can drop parts of multi-requirement prompts; GPT models show more consistent instruction adherence; and stronger models (Opus 4.7, GPT-4.5) write self-verification tests far more frequently than weaker ones — a divergence that emerges specifically because DeepSWE does not instruct models either to write or avoid tests. Essential viewing for teams building coding agents or designing evaluation frameworks.
📺 Source: AI Engineer · Published July 26, 2026
🏷️ Format: Keynote Launch







