Fable 5.1 Just Dropped. It’s Not Even Close.

Fable 5.1 Just Dropped. It’s Not Even Close.

More

Summary

Nick Saraev delivers a fast-paced benchmark breakdown of Anthropic’s Fable 5.1 and Mythos 5.1 releases, covering five major evaluation categories with specific scores and direct comparisons to Fable 5, Opus 5, and GPT-5.6-Sol. The headline number is on Terminal Bench Science 1.0’s Agent Scientific Research task: Fable 5.1 clears 50%, more than doubling the previous generation’s sub-30% ceiling and decisively outperforming GPT-5.6-Sol. Agentic coding reaches 55.8% versus Fable 5’s 42% and GPT-5.6-Sol’s 37.3%. On OSWorld 2.0 computer use, Fable 5.1 scores 77.9% — a notable catch-up in an area where Anthropic’s models had historically trailed OpenAI’s.

The most commercially significant data point comes from Automation Bench, which measures a model’s reliability at automating real business pipelines: Fable 5.1 scores 31.4% versus Fable 5’s 17.1%, nearly doubling the rate at which the model can reliably complete end-to-end workflow automation. Saraev interprets this as a signal that many more businesses will find it economically viable to automate workflows with Fable 5.1 at current API pricing.

A cost-efficiency frontier graph shows approximately 2.5x better performance per dollar at comparable spend levels versus Fable 5. Saraev argues that cost-per-task — not cost-per-token — is becoming the correct unit of measurement as more capable models complete complex work in fewer but more effective steps. Cache read pricing has been cut significantly, with up to 45% savings for heavy agentic workloads. The video closes with a broader note on the gap between benchmark rankings and real-world experience, citing Opus 5 as a model that scored well but underdelivered in practice.


📺 Source: Nick Saraev · Published September 01, 2026
🏷️ Format: News Analysis

1 Item

Channels