Claude Fable 5.1 Is Not More Token Efficient
Claude Fable 5.1 pricing: $10 in, $50 out with a 128K output ceiling, but the real savings come from a 75% cache-read cut, only if the same material repeats within five minutes. Claude Fable 5.1 kept Fable 5's base prices intact, $10 per million input tokens and $50 per million output on OpenRouter and Anthropic's platform, but the only number that moved is the cache-read rate, which dropped 75% from $1 to $0.25 per million tokens. The headline million-token window comes with an uncomfortable ceiling: maximum output is capped at 128,000 tokens, making the real worst-case cost roughly $16.40 per request, not the $60 naive math suggests. The discount is real but conditional. It only applies if the same material re-enters within five minutes (the cheap cache tier), if nothing in the earlier context gets edited (which voids the cache), and if the window actually needs refilling, meeting all three conditions unlocks savings up to 45% for agentic workloads, but miss any one and you've paid a storage premium for a discount you can't collect. Single-pass work like document review or one-off classification jobs hit the $12.50 cache-write fee and never benefit, while fixed-context agent loops running continuously see the full cut. For one-shot tasks, smaller curated inputs beat filling the window, which is why Anthropic's own guidance sends most builders to the cheaper Claude Opus 5 first. A worked example shows fifty agent-loop iterations costing $12 with caching versus $100 cold, illustrating the real money sits in repetition, not in the window size itself. This video breaks down when Fable 5.1 actually saves money and when it's a hidden surcharge, for builders deciding between large windows, caching, and retrieval-based approaches. Chapters: 0:00 The price that didn't move 0:28 Ten bucks fills it all 1:35 Why 128K caps the output 2:54 Where the real savings hide 4:07 The entry fee clock 5:40 Two sends break even 6:53 Tool bloat eats your window 8:21 One edit kills the cache 9:42 One-shot work pays to wait 11:45 Anthropic starts you cheaper 12:59 Agent math: $12 vs $100 14:27 When 75% actually matters 16:23 Is it cheaper or not? Tools & resources mentioned: - OpenRouter: https://openrouter.ai - Anthropic Claude Platform: https://platform.claude.com - Claude Models Overview: https://platform.claude.com/docs/en/models/fable-5-1/overview About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #fable5.1 #llmpricing #claudeai