Summary
Matthew Berman delivers a structured breakdown of Claude Opus 4.8, Anthropic’s latest flagship released approximately six weeks after Opus 4.7. On Swebench Pro, Opus 4.8 scores 69.2% โ a 5-point improvement over 4.7, roughly 10 points above GPT 5.5, and well ahead of Gemini 3.1 Pro. GPT 5.5 retains the lead on Terminal Bench 2.1 for agentic terminal navigation at 78.2%, which Berman cites as a likely driver of the persistent real-world preference for GPT in coding tasks despite Swebench gaps. He also flags the Deepu benchmark as a more reliable proxy for practical performance than Swebench scores alone.
A significant portion covers dynamic workflows, Anthropic’s new Claude Code research preview (available in CLI, desktop, VS Code, and via Amazon Bedrock, Vertex AI, and Microsoft Foundry) that dynamically writes orchestration scripts running tens to hundreds of parallel sub-agents in a single session. Berman connects this to Anthropic’s recently secured compute capacity through a partnership with xAI for access to the Colossus infrastructure โ arguing the company has been sitting on these features and is now releasing them as compute constraints ease.
Pricing for regular usage is unchanged at $5 per million input tokens and $25 per million output tokens. Fast mode is now three times cheaper than before and 2.5x faster. The video also covers the Claude Mythos preview โ a model Anthropic describes as a new intelligence tier above Opus, currently limited to cybersecurity use cases โ with a broader release expected within weeks.
๐บ Source: Matthew Berman ยท Published May 29, 2026
๐ท๏ธ Format: Review







