Is Kimi K3 Really That Good?! (Don't Just Believe The Hype)

From the creator

Kimi K3 is the most powerful open weight model ever released, and the benchmarks make it look like it's beating Opus 4.8 and almost matching the very best closed models. After spending millions of tokens putting it up against Opus 4.8 and Kimi K2.7 on real engineering tasks, I'm not buying it, and I don't think you should either. Kimi K3 is genuinely impressive, and sometimes the raw output is even better than Opus. But open weight models have reliability failure modes that the public benchmarks don't seem to account for. In this video I show you exactly where Kimi K3 breaks down, why it's still useful as the workhorse in a mixed-model workflow, and how to build your own benchmarks that actually reflect real world coding instead of a leaderboard. Everything is open source so you can run these tests yourself, even with other models! ~~~~~~~~~~~~~~~~~~~~~~~~~~ - QA.tech, the AI testing tool with autonomous QA agents that test your app like a real user: http://qa.tech/cole - Past the Bottleneck - free guide to releasing quality products in the AI-SDLC, for teams that ship fast and want to stay safe: http://qa.tech/cole-book ~~~~~~~~~~~~~~~~~~~~~~~~~~ - Join my free live workshop on July 29th, where I show you how to become an AI native engineering organization - building a reliable standard for how teams use AI coding agents: https://dynamous.ai/ai-native-engineering-org - Benchmark repo (workflows, rubric, and every exact prompt): https://github.com/coleam00/kimi-k3-reliability-benchmark - Archon (the open source harness builder I used): https://github.com/coleam00/archon ~~~~~~~~~~~~~~~~~~~~~~~~~~ 0:00 The Kimi K3 Hype vs Reality 1:41 The Benchmark Suite I Built 3:43 Why Open Weight Models Are Worth It 4:47 Mixing Models for a Cheaper Workflow 5:50 Benchmark 1: Real Engineering Tasks 8:16 Sponsor: QA.tech 10:19 How the Scoring Works (7 Dimensions) 10:56 Simple Tasks: K3 Basically Ties Opus 12:28 Complex Tasks: Opus Pulls Ahead 14:17 Why Public Benchmarks Can't Be Trusted 15:33 The Trap Tasks Explained 17:00 The Results: 8% vs 36% Failure Rate 18:58 Why Opus Thinks for Itself 20:12 The Takeaway: Plan Big, Build Cheap ~~~~~~~~~~~~~~~~~~~~~~~~~~ Join me as I push the limits of what is possible with AI. I'll be uploading videos weekly - at least every Wednesday at 7:00 PM CDT!

Choose to Build with AI
Matched to AI Agents

AI Maker Residence 3

The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence 3
Live event
AI Maker Residence 3
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.