Anthropic went CRAZY (Opus 5.5)

Anthropic went CRAZY (Opus 5.5)

More

Summary

Matthew Berman breaks down Anthropic’s launch of Claude Opus 5.5, the company’s first model release since Dario Amodei’s widely discussed essay calling for the industry to “pace the frontier.” Joined by Tharic, a member of Anthropic’s technical staff, Berman walks through benchmark results showing Opus 5.5 topping Claude Fable 5.1 and GPT-6-Astra on Terminal Bench 4.0 (66.4%), GDPval (1846 ELO), and Frontier Code, while costing 40% less to run than Opus 5.

The conversation covers how Anthropic achieved the cost reduction through both lower per-token pricing ($4/$20 per million input/output tokens) and improved compute efficiency, along with a discussion of the model’s dual-use safety testing in biology and cybersecurity, conducted with external evaluators including METER and Frontier Design. Tharic explains the reasoning behind Anthropic’s expanded alignment testing for long-horizon and “impossible” tasks, referencing incidents like the Hugging Face exploit.

Berman also reviews Opus 5.5’s strong showing on Artificial Analysis’s intelligence index, where it takes the top spot with a score of 58, ahead of Fable 5.1 and GPT-6-Astra at 53. The video closes with reflections on how efficiency gains in frontier models like Opus 5.5 signal what capabilities may eventually run locally, alongside a note on rising GPU hardware costs affecting local AI enthusiasts.


📺 Source: Matthew Berman · Published September 23, 2026
🏷️ Format: News Analysis

1 Item

Channels

1 Item

Companies