Sonnet 5 (Fully Tested): IT UNDERPERFORMS GLM-5.2 and COSTS MORE!?
Google Instant Ramen Test Video: https://youtu.be/tarRi_wveUA In this video, I'll be telling you about the new Sonnet 5 model, its benchmark results, pricing, coding performance, and why I think it is not a very good release compared to models like Opus and GLM-5.2. -- Key Takeaways: 🚀 Sonnet 5 is finally here, and it is positioned as a cheaper alternative with performance close to Opus 4.8. 💸 The model launches with introductory pricing of $2 per million input tokens and $10 per million output tokens until August 31, 2026. 📊 Anthropic’s benchmark chart is surprisingly small, with only a few benchmarks like SWE Bench Pro, Terminal Bench, HLE, OS World Verified, and GDP Eval. ⚠️ Sonnet 5 generally scores lower than Opus 4.8 across most benchmarks and evaluations. 🧪 In real-world testing, Sonnet 5 struggled with multiple coding and generation tasks, including ThreeJS, SVG, simulators, and math questions. 💻 The model behaves strangely in coding tools like Claude Code and OpenCode, including working in root temp folders and not following instructions well. 📉 Compared to GLM-5.2, Sonnet 5 costs significantly more while delivering weaker performance in my testing. ❌ Overall, Sonnet 5 feels overpriced, underwhelming, and not very useful compared to better alternatives like Opus or GLM-5.2.