The 80B Model That Fits In 4.3 GB Of RAM
Mixture-of-experts architecture lets 80B parameter Qwen models run on Mac in 4.3 GB RAM via Swiftlet's disk-streaming runtime, and a 35B sibling on iPhone, though with caveats. Swiftlet is a Swift+Metal runtime that streams mixture-of-experts weights from SSD to run Qwen MoE models locally on Apple Silicon with minimal RAM. The 80B-parameter Qwen3-Next-80B-A3B activates only ~3B parameters per token across 512 routed experts, letting Swiftlet keep just the 2.5 GB dense core resident while fetching unused experts from disk on demand, a trick that trades speed for memory. On a base M5 MacBook (24GB unified memory), the full 80B model peaks at 4.3 GB RAM and decodes at 4.5, 5 tokens/second; on an iPhone 17, the 35B sibling (Qwen3.6-35B-A3B) runs at ~2.5 GB RAM and 1 tok/s after a 30-second first reply. The 35B ships in Priv AI via an 18 GB resumable download; the 80B requires 42 GB disk on Mac. Outside testing (a 2020 M1 Mac mini, a 64GB MacBook Pro) confirmed the cache doesn't bottleneck speed, the GPU dispatch is the constraint, and validated that a 397B model can stream at 1.4 tok/s on hardware that cannot hold it resident. The cost: storage (18, 42 GB), latency on long prompts (minutes for 591-token system prompt), and 3B-parameter recall ceiling despite 80B model fluency. For builders and AI-curious engineers evaluating local inference on Apple Silicon versus cloud APIs, this shows the RAM-speed tradeoff and when storage-streaming actually wins. Chapters: 0:00 80 billion parameters, 2.5 gigabytes of RAM 0:23 The M5 MacBook goes toe-to-toe with a server 2:01 Why only ten experts wake up per word 3:26 The hidden cost of loading every expert 5:02 One SSD read per word, forty-two gigabytes total 6:38 Why the disk isn't the bottleneck here 8:03 The iPhone gets the smaller sibling instead 9:33 Eighteen gigabytes lands on your phone today 10:58 A 2020 Mac mini breaks the maintainer's numbers 12:38 Big model talk, small model memory 14:13 When a laptop can't hold the model at all 16:00 A 397 billion parameter model on an iPhone Pro 17:39 What you actually get, machine by machine Tools & resources mentioned: - Swiftlet: https://github.com/leonickson1/Swiftlet - Priv AI: https://privy.app - Qwen3-Next-80B-A3B: https://huggingface.co/Qwen/Qwen3-Next-80B-A3B - MLX: https://github.com/ml-explore/mlx - ANEMLL About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #moe #apple #llm