These AI Models Actually Run On A 96GB Mac

From the creator

Run local AI on 96GB Mac: model sizes that fit, speeds real machines achieve, and why 120B mixture-of-experts beats 70B dense. A 96 GB Mac does not give AI models all 96 GB. By default macOS lets the graphics side use about three quarters of it: about 77 GB in Hugging Face download page numbers (72 GiB), and around 83 GB on newer macOS releases judging by one recent log. Inside that budget, the size that runs well depends on how much of the file each token has to read. About 30B dense: Qwen 3.8 27B is 16.1 GB at 4 bit and 29.5 GB at 8 bit, so even the 8 bit file leaves nearly 50 GB for long chats. A user run on the oMLX benchmark list measured 53.7 tokens a second on the $5,499 base M5 Ultra with 96 GB. Dense 70B fits but is usually the wrong pick: every token reads the whole 39.7 GB Llama 3.3 70B file. A user run on a 64 core M5 Ultra (the base model's graphics chip, in a 256 GB machine) measured 24.5 tokens a second, under half the 27B. At a 64K token prompt memory grew to 57.9 GB, writing slowed to 14.8 and the first word took about three minutes. About 120B mixture of experts is the largest normal fit and the fastest, because each token reads only a slice: gpt-oss-120b (117B parameters, 5.1B active, 63.4 GB) measured 104.3 tokens a second on the same 64 core chip. For Qwen 3.5 122B, pick the 69.6 GB MLX or 71.7 GB GGUF file, not the 76.5 GB one. Qwen 3.8 Flash Next (106 to 112 GB) fits only in oMLX, by parking its roughly 30 GB lookup table on the SSD; fresh prompts get slower. The 200B class does not fit: DeepSeek V4 Flash is about 155 GB at 4 bit and GLM 5.3 Flash is still 93.1 GB at 1 bit. Tom's Hardware lists 96 GB to 256 GB as a $4,000 upgrade. Verdict: run a 120B mixture of experts in its smaller 4 bit download; for 8 bit or maximum headroom, a dense model around 30B; skip dense 70B. Check your own budget first (LM Studio has shown it as VRAM). Speeds are their owners' and reviewers' published results; The Stack did not run these models. Sources and media credits: Benchmarks: oMLX public benchmark list (user submissions) and oMLX issue #4140; MacStories (Federico Viticci), Tom's Hardware and BGR reviews Memory budget: Apple Metal documentation, Ollama issue #10626, drama_llama issues #125 and #137, LM Studio, llama.cpp discussion #2182, mlx-lm Model files: Hugging Face pages of Qwen, OpenAI, mlx-community, lmstudio-community and unsloth llama.cpp marks: github.com/ggml-org/llama.brand Stock footage: Tima Miroshnichenko and Kaboompics (Pexels) Chapters: 0:00 Intro 2:06 Real limit 4:33 Small models 8:22 70B models 11:28 120B models 15:32 Flash Next 18:46 Too big 20:41 Conclusion Tools & resources mentioned: - LM Studio: https://lmstudio.ai - Ollama: https://ollama.com - oMLX: https://omlx.ai - mlx-lm: https://github.com/ml-explore/mlx-lm - llama.cpp: https://github.com/ggml-org/llama.cpp - Qwen 3.8 27B (4 bit MLX): https://huggingface.co/lmstudio-community/Qwen3.8-27B-MLX-4bit - gpt-oss-120b: https://huggingface.co/openai/gpt-oss-120b - Qwen 3.5 122B: https://huggingface.co/Qwen/Qwen3.5-122B-A10B - Qwen 3.8 Flash Next: https://huggingface.co/Qwen/Qwen3.8-Flash-Next - oMLX SSD offload report (#4140): https://github.com/jundot/omlx/issues/4140 - MacStories M5 Ultra review: https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/ About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #local ai #llm #mac studio

Choose to Build with AI
Matched to Paged Attention

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.