Fable 5.1 Just Dropped: It’s 45% Cheaper & 2X Better

From the creator

https://bitbiased.ai/ai-automation-services Anthropic says Fable 5.1 can be 45% cheaper while delivering more than double the performance of Fable 5 on some benchmarks. Both claims are technically real — but independent testing reveals a much stranger story. Fable 5.1 arrived on September 1st, less than three months after Fable 5, and Anthropic is positioning the point upgrade as its new flagship for long, multi-step coding, scientific reasoning, research, and agentic workflows. The underlying specs haven't dramatically changed: the model keeps the one-million-token context window, 128,000-token output cap, and the same $10 per million input tokens and $50 per million output tokens. What has changed is how the model uses those resources — and one setting can completely change whether Fable 5.1 looks like an incredible value or an expensive overkill machine. The biggest performance gains are concentrated in long-horizon tasks. On Terminal-Bench-Science, Fable 5.1 scored 52.6%, compared with just 24.7% for Fable 5. AutomationBench jumped from 17.1% to 31.4%. But on shorter general-reasoning benchmarks, the difference is far less dramatic: Humanity's Last Exam moved from 57.8% to 60.9% without tools, while GDPval-AA reached 1,853 Elo versus Fable 5's 1,723. Then there's the pricing story. Anthropic cut prompt-cache read pricing by 75%, from $1 per thousand cached tokens to $0.25. For agentic workloads that repeatedly reuse large amounts of context, Anthropic says that can translate into roughly 25% lower typical costs and savings of up to 45%. But independent testing from Artificial Analysis found something the headline pricing doesn't capture: at maximum effort, Fable 5.1 reportedly used roughly 1.7x as many output tokens as Fable 5 on the same composite task. That pushed the measured cost to $3.76 per task — around 20% more than Fable 5. Even at a reduced "extra-high" effort setting, the measured $2.72 cost remained above Opus 5's $2.34. At low effort, however, the picture flips. Testing suggests Fable 5.1 can match Fable 5's quality for roughly one-quarter of the cost. The effort setting isn't just a minor optimization anymore — it can effectively turn Fable 5.1 into two very different products. There's another unusual detail: Fable 5.1 and Mythos 5.1 reportedly use the exact same underlying weights. The difference is in their safety filters. Fable is broadly available to paid users, while Mythos is restricted to vetted organizations working in areas such as cybersecurity research and life sciences. That distinction actually appears in benchmark results. Fable 5.1 scored 55.8% on Terminal-Bench 4.0, while Mythos 5.1 reached 60.9%. Anthropic's methodology indicates that safety refusals can count as zeroes on certain benchmarks, meaning some of the apparent performance difference comes from what each version is permitted to attempt rather than a difference in the underlying model. Independent testing also gives Anthropic's broader performance claims some credibility. Artificial Analysis measured an Intelligence Index score of 66 for Fable 5.1 at maximum effort, ahead of Opus 5 at 63 and GPT-5.6 Sol at 61. But AA-Omniscience exposed another wrinkle: the model attempted more answers than Fable 5 while also making slightly more mistakes, producing essentially no improvement on that particular test. And there is a privacy catch. Anthropic's new Enterprise Frontier Safeguards can allow qualifying enterprise customers to keep conversation logs on their own cloud infrastructure under customer-controlled keys, with some large customers reportedly eligible for zero data retention. But for regular users, the default 30-day retention policy remains in place. So who is Fable 5.1 actually for? If you're running coding agents across large codebases, long research pipelines, scientific reasoning workloads, or tool-heavy tasks that repeatedly reuse context, the upgrade could be substantial. If you're mainly doing quick Q&A, light writing, or short prompts, the performance gains are much smaller — while Fable 5.1's pricing and latency can make Opus 5 or Sonnet 5 more sensible choices. The real Fable 5.1 story isn't simply that Anthropic made a model that's "45% cheaper" or "twice as good." It's that the model's value changes dramatically depending on workload, cache reuse, and especially the effort level you select. And that makes the effort slider arguably more important than the model name itself. #anthropic #fable51 #ai #artificialintelligence #llm

Choose to Build with AI
Matched to AI Developer

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.