Apple's New M5 Ultra & M6 Runs Huge Local AI Models
The $5,499 M5 Ultra Mac Studio fits a 122B-parameter LLM, and reviews of the top chip show it writing as fast as a cloud rental, but it only beats cloud pricing for nonstop agents on private data. Apple's M5 Ultra Mac Studio went on sale September 22 at $5,499 with 96GB unified memory and claims to run huge local LLMs entirely on device. The 96GB entry model fits Alibaba's Qwen 122B once it is squeezed to 4-bit (a 71.7GB file). Reviewers tested the top-tier chip with 256GB, priced above $10,000, and it writes roughly 80 tokens per second, matching an ordinary cloud rental of the same model, while Cerebras ran a smaller Qwen model at about 1,800 tokens/sec. Nobody has published a test of the $5,499 version yet. The 50% faster memory bandwidth (1.2TB/s vs M3 Ultra's 819GB/s) drives most of the speed gain; the M5 Max Mac Studio at $5,399 offers 128GB, but its memory reads at about half the speed (614GB/s). Cost math flips based on workload: for everyday chat, $5,499 buys about 23 years of Anthropic's $20/month Claude Pro plan, but always-on agents change the math, MacStories' Federico Viticci ran his agents around the clock for 99 days on an older Mac Studio at no cost beyond the computer and electricity, where paying per token would have been cost-prohibitive. The base model cannot hold the bigger open models (DeepSeek V4 Flash is 137, 155GB even at 4-bit, which means the $9,499 256GB version), and the smartest models from OpenAI, Anthropic and Google can't be downloaded at any memory size. The memory is soldered in, so there is no upgrading later, and on September 24 Apple's US store quoted 7, 8 weeks to ship the base Ultra. This machine is for builders running private-data agents nonstop; for chatbot use or the smartest models, cloud remains the sane choice. Chapters: 0:00 Intro 1:47 What fits 3:41 Speed 6:02 Cost 8:25 Worth it? 10:22 Conclusion Tools & resources mentioned: - Qwen (Alibaba) - OpenRouter - Cerebras - llama.cpp - DeepSeek V4 Flash About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #mac studio m5 ultra #local llm