Kimi K3: An Open Source #1 — Coding Tests vs Fable 5, Opus 4.8 & GPT-5.6 Sol
Kimi K3 is Moonshot AI's new 2.8-trillion-parameter open model — the first open model ever in the 3T class, with a 1M-token context window and native multimodal input (text, image and video). Demand was so high at launch that Kimi had to pause new subscriptions. In this video, I break down why it's such a big deal: it's the first open model to top the Design Arena frontend coding leaderboard, it sits at the top of the Artificial Analysis Intelligence Index among open models, and it goes toe-to-toe with Claude Fable 5 and GPT-5.6 Sol on Terminal-Bench 2.1. I also run my own benchmark on a real GitHub issue — comparing quality, cost, time per task and token usage vs Claude Fable 5, Claude Opus 4.8 and GPT-5.6 Sol — and show you how to run the same evals yourself with DuoBench, plus how to start using K3 today in Pi, Tau and OpenCode. --- 🔗 *Links* - Kimi K3 official blog: https://www.kimi.com/blog/kimi-k3 - Kimi K3 quickstart docs: https://platform.kimi.ai/docs/guide/kimi-k3-quickstart - Kimi K3 on Artificial Analysis: https://artificialanalysis.ai/models/kimi-k3 - Design Arena Code WebDev leaderboard: https://arena.ai/leaderboard/code/webdev/overall - Kimi's announcement about pausing new subscriptions: https://x.com/Kimi_Moonshot/status/2078855608565207130 - DuoBench — run the same evals on your own repos: https://github.com/alejandro-ao/duobench - Tau: https://github.com/huggingface/tau - Pi: https://github.com/earendil-works/pi - OpenCode: https://github.com/anomalyco/opencode --- 👋 *Connect with me* - My website: https://alejandro-ao.com/ - X (Twitter): https://x.com/_alejandroao - LinkedIn: https://www.linkedin.com/in/alejandro-ao/ --- 🤓 *Topics Covered* - Kimi K3 benchmarks vs Claude Fable 5, Opus 4.8 & GPT-5.6 Sol - Testing coding models on real GitHub issues with DuoBench - How to use Kimi K3 in Pi, Tau & OpenCode --- ⏱️ *Timestamps* 0:00 Introduction: Kimi K3 is a DeepSeek moment 1:35 Specs: 2.8T params, 1M context, multimodal, open weights July 27 4:11 Design Arena: first open model #1 in frontend coding 6:39 Artificial Analysis Intelligence Index 7:42 Terminal-Bench 2.1 & token usage 10:24 My own tests: quality vs cost vs time per task 14:43 Run your own evals with DuoBench 16:04 Using Kimi K3 in Pi, Tau & OpenCode 18:23 Conclusion