SWE-2 (Fully Tested): WHAT? IT ACTUALLY BEATS ASTRA & FABLE!
In this video, I’ll be testing Cognition’s new SWE 2 coding model to see how well it performs on KingBench 3, longer agentic tasks, interactive simulations, 3D applications, games, and model fine-tuning. I’ll also compare SWE 2 with DeepSeek V4.1 Flash, discuss its clarification problem, and explain why its inclusion in the $20 Devin Pro plan makes it an interesting AI coding option. -- Key Takeaways: 🚀 SWE 2 scored 67 out of 80, or 83.75%, on KingBench 3. ⚔️ It narrowly beat DeepSeek V4.1 Flash, which scored 81.25%. 🏆 SWE 2 currently ranks fifth on the KingBench 3 leaderboard. 🧠 The model performed especially well on longer, connected tasks involving fine-tuning and local web applications. 📈 SWE 2 finished ahead of Kimi K3, Opus 5, and Fable 5 in this benchmark. ❓ Its biggest drawback is that it often asks too many clarification questions before starting the work. 💸 SWE 2 is currently included at no additional cost with the $20 Devin Pro plan through October 10, 2026. 👍 Overall, SWE 2 is a capable and consistent AI coding model that is well worth testing on real projects.