Xiaomi's AI Cube Vs NVIDIA DGX Spark

From the creator

Xiaomi's AI Cube vs DGX Spark: why bandwidth specs split across two chips, no benchmarks exist, and the prototype won't ship until 2027. Xiaomi's AI Cube is being called four and a half times faster than NVIDIA's DGX Spark, but the 1.22 TB/s bandwidth claim belongs to the O100 chip while the 160GB memory belongs to the separate D100 chip, so no single prototype matches the viral specs. This video walks the actual comparison: memory bandwidth governs local LLM decode speed at batch size 1, and the O100's throughput would be genuinely impressive if it could be independently verified. But DGX Spark is a shipping product ($3,999, $4,699 since October 2025) with published Ollama benchmarks showing 41, 58 tok/s on open-weight models, while the AI Cube has zero independent third-party benchmarks, every spec comes from Xiaomi's own presentation. Even a fair test would require matched quantization, identical inference runtime (vLLM vs Ollama change Spark's throughput by 2.7x on the same silicon), and production availability. Xiaomi projects commercial availability in 2027, the O100/D100 chips are built on experimental wafer-on-wafer packaging (uncertain manufacturability without EUV access), and no price exists yet. NVIDIA's box is real; the prototype is a demonstration. For builders choosing local inference hardware today, DGX Spark has a known ceiling and a supply chain. The Cube is a fascinating hardware direction, but you can't deploy what you can't buy. Chapters: 0:00 One bandwidth, two separate chips 1:35 Bandwidth starves everything else 3:07 Why the Spark costs money 4:25 Even NVIDIA's machine has limits 5:49 Xiaomi's numbers live in marketing 7:00 Software changes speed more than silicon 8:26 The viral claim got debunked instantly 9:49 The packaging problem nobody mentions 11:21 The box has no checkout button 12:45 Meet it in twenty twenty seven 14:22 What settling this actually requires 15:34 Replace nothing you can't yet order Tools & resources mentioned: - NVIDIA DGX Spark: https://docs.nvidia.com/dgx/dgx-spark/hardware.html - Ollama: https://ollama.ai - vLLM - NYU Shanghai RITS Brief - VideoCardz About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #xiaomi ai cube #dgx spark #local llm inference

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.