Kimi K3 Can Run Locally With Exo | Decentralised AI
Kimi K3 needs 16 B200 GPUs to run. exo clusters your devices to bypass that floor, but does it actually work on hardware you own? Kimi K3, Moonshot AI's 2.8-trillion-parameter open-weight model released July 2026, demands sixteen NVIDIA B200 GPUs or eight B300s to run at native precision, a datacenter-scale floor, not a desktop. exo is a disaggregated-inference framework that splits models across your machines using tensor parallelism and RDMA over Thunderbolt 5, claiming a 3.2x speedup on four devices. The pitch sounds compelling: match compute-heavy prefill to an NVIDIA DGX Spark and memory-bandwidth-heavy decode to an Apple Mac Studio, stream the KV cache between them, and get the best of both worlds. Testing on Llama 3.1 8B, exo's official benchmark showed a 2.8x speedup (6.42s → 2.32s), but that entire gain lives in prefill on a 32-token generation task. When decode ran in isolation, the Mac Studio's speed was identical with or without the Spark, the clustering only made you wait less, not generate faster. Scaling to real frontier models exposes the limits: a 128GB DGX Spark hits ~2 tokens/sec on Kimi K3 with Q4_K_M quantization (97s prefill); a community streaming project on an M1 MacBook reaches 4.1 tokens/minute; sixteen H200s running SGLang hit 16.8 tokens/sec; and the only public reproduction attempt fell back to CPU. Meanwhile, OpenRouter hosts Kimi K3 across nineteen providers at wildly different speeds (6, 92 tokens/sec), pricing ($2.40, $6/1M tokens), and precision (some substitute FP8 for the native MXFP4). The real win: long-context prompts with short answers, where prefill dominates and the network transfer hides behind compute. For anyone running thousand-token codebases seeking surgical fixes, the math flips. For everyone else, local clustering remains unproven on consumer hardware, and the cloud is a lottery by provider, not a single product. This is for builders evaluating whether to cluster their own machines, route through OpenRouter, or chase the unverified local frontier. Chapters: 0:00 Turning Four Boxes Into One Brain 1:36 Prefill vs. Decode: Why They Need Different Hardware 3:34 The Odd Couple: DGX Meets Mac Studio 5:35 When 6.4 Seconds Becomes 2.3 6:57 DeepSeek at 671B on Your Desk 8:07 Why 32 Tokens Breaks the Speedup Myth 9:16 Where Did Those Gains Actually Hide? 11:13 Kimi K3's 350x Scaling Problem 12:21 The $500K Baseline Before Compression 13:54 Two Tokens Per Second on Enterprise Silicon 15:18 Four Tokens a Minute on Apple Silicon 16:39 When Even 16 H200s Can't Load It 17:40 The Demo That Fell Back to CPU 19:22 Nineteen Products Behind One Model Name 21:45 The One Shape That Makes Clustering Win 22:52 Local, Frontier, or Just Route Through OpenRouter? Tools & resources mentioned: - exo: https://github.com/exo-explore/exo - OpenRouter: https://openrouter.ai/moonshotai/kimi-k3:batch About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #openrouter #localai #kimi k3