Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs

Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs

More

Descriptions:

Denys Linkov, ML lead at Wisedocs, presents findings from a six-month production refactor of the company’s AI pipeline at the AI Engineer conference. Wisedocs processes complex medical claims in PDFs exceeding 10,000 pages, and its legacy codebase — spread across more than 10 repositories — had grown too slow and too painful to maintain, leading the team to undertake a full architectural overhaul starting in April 2025.

The talk benchmarks AI coding agents specifically on legacy codebases versus clean greenfield code, documenting the real-world performance gap teams face when applying tools like Claude Code to inherited systems rather than new projects. The team evaluated five open-source orchestration frameworks against 17 criteria and built proof-of-concepts with three engineers — a process Linkov estimates could now run 90% faster with modern deep research and agentic tooling. Before-and-after comparisons show meaningful gains in both shipping velocity and code maintainability once the refactor was complete.

A central argument in the talk concerns accuracy thresholds for agentic workflows: while the Metr benchmark graph is commonly shown at 50% task completion, Linkov argues that 80–99% accuracy is the only operationally meaningful target. At 50%, a one-hour agent run has a coin-flip chance of producing nothing useful — a poor trade-off when engineering attention is the scarcest resource. The talk draws on Anthropic case studies from Spotify and Stripe alongside Wisedocs’ own production data to ground these claims.


📺 Source: AI Engineer · Published August 08, 2026
🏷️ Format: Workflow Case Study

1 Item

Channels