Descriptions:
Matthew Berman walks through the Claude Opus 5 release from Anthropic, systematically covering benchmark results across agentic coding, computer use, knowledge work, and reasoning evaluations. The headline finding: Opus 5 outperforms Claude Fable 5 on most benchmarks while pricing at parity with its predecessor Opus 4.8 — a result Berman describes as genuinely unexpected given Fable 5’s position as Anthropic’s flagship model.
Standout numbers include a jump to 43% on agentic terminal coding (versus Fable 5’s 33%), a 30% score on ARC-AGI 3 where previous frontier models topped out near 8%, a four-point improvement on OS World computer use, and a nine-point gain on Automation Bench. The video also includes sponsored benchmark data from Box, which tested Opus 5 on realistic enterprise document tasks across 12 industries — document review improved from 63% to 78% and due diligence from 65% to 76% versus Opus 4.8, with the largest gains concentrated on multi-step analytical work.
Berman frames cost-per-task as the metric that matters most this generation, noting that raw per-token pricing is misleading when models consume very different token counts to complete identical tasks. He identifies a slight regression on the Deep Suite benchmark and intentional underperformance on cybersecurity exploit tasks as the two meaningful caveats, and speculates that Anthropic used Fable 5’s architecture to optimize Opus 5 into a more efficient package — likely the model most practitioners should default to for general knowledge work and coding automation.
📺 Source: Matthew Berman · Published July 24, 2026
🏷️ Format: News Analysis







