The New Mac Studio Makes Local AI 4X Faster
The M5 Ultra's clearest local AI gain is the wait before the first word: 13.9 seconds down to 5.6 in MacStories' test, while text generation rose about 54%. In a matched comparison by Federico Viticci at MacStories (Qwen3.8 Flash-Next, identical software), a fresh prompt of roughly 16,000 tokens took 13.9 seconds to produce the first word on an M3 Ultra and 5.6 seconds on the M5 Ultra. That test put a 256 GB M5 Ultra against a 512 GB M3 Ultra, not base configurations, and it measures only the initial delay, not total completion time or output quality. In separate MacStories throughput testing, text generation rose from 70 to 108 tokens per second, about 54%: reading your prompt (prefill) is parallel math, while writing the answer one token at a time (decode) is usually limited by memory bandwidth, as NVIDIA's inference guide explains. Apple put a Neural Accelerator in every GPU core and quotes up to 4.5x the peak GPU compute for AI, a theoretical ceiling rather than app speed, plus 1.2 TB/s of memory bandwidth, 50% more than the M3 Ultra. No single multiplier describes both jobs, and software needs explicit support: current MLX versions route that matrix math to the new hardware. Apple lists 96 GB to start, with 256 GB and 512 GB on the top chip, and the Apple Store says the 512 GB option is coming late October. If a model already fits on a graphics card, a PC can still be faster: in Viticci's whole-system comparison an RTX 5090 with an 8-bit attention cache generated text faster at the longest context. Before buying, check your software: prompt caching (in MLX, for example) reuses an already processed document for follow-up questions until an edit near the top forces a re-read. Then run a fair test on the machine you own, timing both the first word and a complete, useful answer. Image, transcription, vision, agent and video jobs are different workloads; Apple shows Draw Things image generation on the M5 Ultra, but our research found no matched speed comparisons for them. The M5 Ultra Mac Studio starts at $5,499 in the US with 96 GB. The verdict: if your current setup handles your daily work, keep it; consider the M5 Ultra when repeated long-document delays or a lack of memory for local models are real constraints. Media credits: Footage: Apple, "The New Mac Studio with M5 Max and M5 Ultra" https://www.youtube.com/watch?v=3uAIqqg8ZHo Stock footage (Pexels): Towfiqu barbhuiya, The MoonRunners Product images: Apple (Mac Studio, M5 Ultra, M3 Ultra, Draw Things on Mac Studio), NVIDIA (GeForce RTX 5090) Test results: Federico Viticci, MacStories Chapters: 0:00 Intro 1:04 The wait 2:43 Hardware 4:40 Memory 6:32 Software 7:57 Other tasks 9:08 Worth it 10:51 Conclusion Tools & resources mentioned: - MacStories: M5 Ultra Mac Studio review by Federico Viticci (the tests): https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/ - Apple Newsroom: M6 and M5 Ultra announcement: https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-performance-and-ai-compute/ - Mac Studio technical specifications (Apple): https://www.apple.com/mac-studio/specs/ - Buy Mac Studio, US prices (Apple Store): https://www.apple.com/shop/buy-mac/mac-studio - Exploring LLMs with MLX and the Neural Accelerators in the M5 GPU (Apple Machine Learning Research): https://machinelearning.apple.com/research/exploring-llms-mlx-m5 - MLX: https://github.com/ml-explore/mlx - MLX LM (prompt caching): https://github.com/ml-explore/mlx-lm - NVIDIA: Mastering LLM Techniques: Inference Optimization: https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/ - Draw Things (App Store): https://apps.apple.com/us/app/draw-things-offline-ai-art/id6444050820 - Qwen3.8 Flash-Next (the model in the first-word test) About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #LocalAI #M5Ultra #MacStudio