These AI Models Actually Run On A 96GB Mac
Run local AI on 96GB Mac: model sizes that fit, speeds real machines achieve, and why 120B mixture-of-experts beats 70B dense. A 96 GB Mac does not give AI models all 96 GB. By default macOS lets the graphics side use about three quarters of it: about 77 GB in Hugging Face download page numbers (72 GiB), and around 83 GB on newer macOS releases judging by one recent log. Inside that budget, the size that runs well depends on how much of the file each token has to read. About 30B dense: Qwen 3.8 27B is 16.1 GB at 4 bit and 29.5 GB at 8 bit, so even the 8 bit file leaves nearly 50 GB for long chats. A user run on the oMLX benchmark list measured 53.7 tokens a second on the $5,499 base M5 Ultra with 96 GB. Dense 70B fits but is usually the wrong pick: every token reads the whole 39.7 GB Llama 3.3 70B file. A user run on a 64 core M5 Ultra (the base model's graphics chip, in a 256 GB machine) measured 24.5 tokens a second, under half the 27B. At a 64K token prompt memory grew to 57.9 GB, writing slowed to 14.8 and the first word took about three minutes. About 120B mixture of experts is the largest normal fit and the fastest, because each token reads only a slice: gpt-oss-120b (117B parameters, 5.1B active, 63.4 GB) measured 104.3 tokens a second on the same 64 core chip. For Qwen 3.5 122B, pick the 69.6 GB MLX or 71.7 GB GGUF file, not the 76.5 GB one. Qwen 3.8 Flash Next (106 to 112 GB) fits only in oMLX, by parking its roughly 30 GB lookup table on the SSD; fresh prompts get slower. The 200B class does not fit: DeepSeek V4 Flash is about 155 GB at 4 bit and GLM 5.3 Flash is still 93.1 GB at 1 bit. Tom's Hardware lists 96 GB to 256 GB as a $4,000 upgrade. Verdict: run a 120B mixture of experts in its smaller 4 bit download; for 8 bit or maximum headroom, a dense model around 30B; skip dense 70B. Check your own budget first (LM Studio has shown it as VRAM). Speeds are their owners' and reviewers' published results; The Stack did not run these models. Sources and media credits: Benchmarks: oMLX public benchmark list (user submissions) and oMLX issue #4140; MacStories (Federico Viticci), Tom's Hardware and BGR reviews Memory budget: Apple Metal documentation, Ollama issue #10626, drama_llama issues #125 and #137, LM Studio, llama.cpp discussion #2182, mlx-lm Model files: Hugging Face pages of Qwen, OpenAI, mlx-community, lmstudio-community and unsloth llama.cpp marks: github.com/ggml-org/llama.brand Stock footage: Tima Miroshnichenko and Kaboompics (Pexels) Chapters: 0:00 Intro 2:06 Real limit 4:33 Small models 8:22 70B models 11:28 120B models 15:32 Flash Next 18:46 Too big 20:41 Conclusion Tools & resources mentioned: - LM Studio: https://lmstudio.ai - Ollama: https://ollama.com - oMLX: https://omlx.ai - mlx-lm: https://github.com/ml-explore/mlx-lm - llama.cpp: https://github.com/ggml-org/llama.cpp - Qwen 3.8 27B (4 bit MLX): https://huggingface.co/lmstudio-community/Qwen3.8-27B-MLX-4bit - gpt-oss-120b: https://huggingface.co/openai/gpt-oss-120b - Qwen 3.5 122B: https://huggingface.co/Qwen/Qwen3.5-122B-A10B - Qwen 3.8 Flash Next: https://huggingface.co/Qwen/Qwen3.8-Flash-Next - oMLX SSD offload report (#4140): https://github.com/jundot/omlx/issues/4140 - MacStories M5 Ultra review: https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/ About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #local ai #llm #mac studio