Mac Studio M5 Ultra Vs Xiaomi's AI Cube
Mac Studio M5 Ultra vs Xiaomi AI Cube for local AI: what fits in 96GB vs 80GB, what actually decides speed, where the software stands, and whether either box is worth buying. Apple's base Mac Studio M5 Ultra ($5,499) packs 96GB unified memory, a 30-core CPU, 64-core GPU, and 1.2TB/s of rated memory bandwidth (a hardware rating, not a measured model speed), and you can order it now. Xiaomi's AI Cube remains an unshipped prototype with 80GB memory and three custom Xring chips, claiming to hold a 120B-parameter model but offering no confirmed price, shipping date, or independent benchmarks. Both machines use unified memory, making model compression and working memory the practical limits. A 4-bit-quantized 120B model needs roughly 60GB before overhead, and the conversation cache needs room on top, so fitting is plausible but not guaranteed for every 120B model. The reviewed sources contain no verified, matched speed test of the base 96GB Mac Studio against Xiaomi's prototype; a fair test would run the same model at the same compression and time both waits, prompt processing and generation. Apple claims up to 4x faster prompt processing in LM Studio than the M3 Ultra, but that is a selective vendor claim, not a measured test on the base 96GB machine, and it says nothing about generation speed or Xiaomi's prototype. Software matters: on the Mac, apps like LM Studio document a working local path today (download a model, load it, chat offline), though a coding agent still needs tool support you should confirm; Xiaomi's runtime, model formats, and developer tools remain unannounced. Memory bandwidth alone does not settle speed: the model architecture and the software also shape how quickly a useful answer arrives. The Mac is worth buying when larger capacity genuinely improves your work, so first check whether a smaller open model on your current computer already handles your daily tasks; the Cube only becomes a purchase option once it ships with a price and verified software compatibility. For builders evaluating local AI workstations, Mac Studio is concrete today; Xiaomi is interesting later. Chapters: 0:00 Intro 1:05 Memory 3:15 Speed 5:21 Software 7:11 Worth it 8:13 Conclusion Tools & resources mentioned: - LM Studio: https://lmstudio.ai/docs/app - llama.cpp - Mac Studio M5 Ultra: https://www.apple.com/mac-studio/ - Mac Studio technical specifications: https://www.apple.com/mac-studio/specs/ - Apple's M5 Ultra announcement: https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/ - Lu Weibing's AI Cube post (Weibo): https://weibo.com/2/detail/5335807696569782 - LM Studio offline operation: https://lmstudio.ai/docs/app/offline About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #MacStudioM5Ultra #LocalAI #LLMs