Opus 5 (Fully Tested): A MID-MODEL for a BIG PRICE that still UNDERPERFORMS K3?!
In this video, I'll be testing the newly released Claude Opus 5 from Anthropic and comparing it against models like Fable 5, Opus 4.8, Qwen 3.8 Max, Kimi K3, and GPT-5.6 Sol using my KingBench benchmark. -- Key Takeaways: 🚀 Claude Opus 5 is Anthropic’s new Opus flagship, priced the same as Opus 4.8 at $5 input and $25 output per million tokens. ⚡ A new Fast mode runs around 2.5 times faster, but costs double, with effort settings for balancing quality and price. 📊 Official benchmarks show huge gains, with Opus 5 beating Fable 5 and Opus 4.8 on Frontier-Bench and ARC-AGI 3. 🧠 On reasoning, math, coding logic, and long-horizon agentic work, Opus 5 performs extremely well and scores several perfect 10s. 🎮 It nails tasks like the elevator simulation, bow and arrow game, math problem, and the local panda finetuning project. 🎨 Visual and 3D tasks are weaker, with regressions on the folding table test and only average results on the panda SVG and wristwatch task. 📉 On my KingBench, Claude Opus 5 scores 62 out of 80, landing below Opus 4.8, Fable 5, and Qwen 3.8 Max. 💬 In real-world use, Opus 5 feels verbose, touches too many files, and seems less capable as a general assistant than Fable 5 or GPT-5.6 Sol. 👍 Overall, Claude Opus 5 is a strong agentic coding and reasoning model, but not a clear upgrade over Opus 4.8 or Fable 5 for every use case.