Sonnet 5.5 Explained: Is it As Good As They Say?

From the creator

Claude Sonnet 5.5 costs less per token than Opus 5.5, but on Artificial Analysis's benchmark tasks at maximum effort its bill per task came out higher: about $7.60 against about $6. Anthropic reports that Sonnet 5.5 delivers about 30% faster output and up to 30% lower task cost than Sonnet 5, at $2 / $10 per million input / output tokens against $4 / $20 for Opus 5.5. It is rolling out on paid GitHub Copilot plans, where organization policy allows it, and is available on OpenRouter. A lower rate does not guarantee a cheaper finished task, because the bill counts everything the model produces, including reasoning tokens. At maximum effort Sonnet 5.5 generated roughly 193,000 output tokens per Intelligence Index task, the most Artificial Analysis says it has measured. These are early figures from pre-release testing: a structured-output bug has since been fixed, and full reruns were pending at launch. At high effort Sonnet was cheaper again, $1.08 against $1.82 per task, but scored lower on the Intelligence Index, 47 against 54 (Artificial Analysis's table as captured on October 1, 2026; it has since been updated). On code, the two CodeRabbit tests favored different models. On 13 difficult known-bug cases, Sonnet 5.5 caught 6 actionable issues, while Opus 5.5 caught 8 in CodeRabbit's Standard review pipeline and 10 in its Max pipeline, a small sample. In one Brick Studio build run, Sonnet finished in 29 minutes 27 seconds against 44 minutes 50 seconds for Opus, with results judged close and Opus slightly higher in fidelity. On office work, Artificial Analysis rates them almost level on GDPval-AA (1844 vs 1846 Elo). On AA-Omniscience, Opus 5.5 was more accurate (66% vs 54%), while Sonnet 5.5 had the lower hallucination rate (47% vs 59%), so neither earns blind trust. On effort, start at the product default (Medium in Claude apps and Claude Code, High on the Claude Platform) and adjust against checked results. More effort buys steps and checks, not knowledge: on CodeRabbit's 80-case set, its Opus 5.5 Standard pipeline caught 51 actionable items to 50 for Max. Updating API code? Sonnet 5.5 rejects requests that explicitly turn thinking off. The verdict: make Sonnet 5.5 your default for bounded, everyday tasks you can verify, and keep Opus 5.5 for intricate multi-file debugging or subtle logic where Sonnet keeps missing requirements. Run one familiar job at the default effort and check both the result and the cost. Media credits: Footage: Anthropic, "Introducing Claude Sonnet 5.5" https://www.youtube.com/watch?v=s5nkj-L2vAw Footage: CodeRabbit, "Claude Sonnet 5.5 vs Opus 5.5: Same Prompt, Side by Side in Claude Code" https://www.youtube.com/watch?v=dUEhC0Ncpqk Stock footage (Pexels): K, Jakub Bukowski Image: GitHub Changelog (Copilot model picker) Test results: Artificial Analysis, CodeRabbit Chapters: 0:00 Intro 1:17 Cost 3:08 Coding 4:40 Office work 6:28 Effort 7:55 Access 9:01 Conclusion Tools & resources mentioned: - Anthropic: Introducing Claude Sonnet 5.5: https://www.anthropic.com/claude-sonnet-5-5 - Artificial Analysis: Claude Sonnet 5.5 launch analysis: https://artificialanalysis.ai/articles/claude-sonnet-5-5 - Artificial Analysis: Sonnet 5.5 vs Opus 5.5 release comparison: https://artificialanalysis.ai/models/releases/comparisons/claude-sonnet-5-5-vs-claude-opus-5-5 - CodeRabbit: Sonnet 5.5 model review: https://www.coderabbit.ai/blog/sonnet-5-5-model-review - CodeRabbit: Opus 5.5 model review: https://www.coderabbit.ai/blog/opus-5-5-model-review - Claude Academy: Choosing the right effort level in Claude Code: https://academy.claude.com/tutorials/choosing-the-right-effort-level-in-claude-code - GitHub Changelog: Claude Sonnet 5.5 in GitHub Copilot: https://github.blog/changelog/2026-09-28-claude-sonnet-5-5-in-github-copilot/ - OpenRouter: Claude Sonnet 5.5: https://openrouter.ai/anthropic/claude-sonnet-5.5 - Claude Platform docs: Migrating to Claude Sonnet 5.5: https://platform.claude.com/docs/en/models/sonnet-5-5/migration-guide About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #claude #ClaudeOpus #AImodels

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.