A Mac Mini Runs A 35B Model In 3GB Of RAM
35B model on Mac via SSD: Edge0 streams mixture-of-experts from disk with 3GB active memory. How it trades memory for bandwidth, and when it breaks. Edge0 is an open-source streaming mixture-of-experts framework that runs Alibaba's 35-billion-parameter Qwen3.5-MoE model on Apple Silicon with under 3 gigabytes of active memory by keeping 92.9% of expert weights on SSD and pulling only the 4 active experts per token into RAM. The 19.51GB checkpoint arrives quantized to 4-bit with trained LoRA recovery adapters and a 33-layer prerouter prediction head that overlaps expert fetches with GPU compute, achieving 15 tokens/sec on a Mac mini M4 Pro. The trade is real: this is a bandwidth problem masquerading as a memory solution. 1.3GB of non-expert weights stay resident; 64 expert slots cache 108MB; only 10 of 40 layers use full attention (the rest are linear attention with no KV cache), so context scales cheaply to 32k tokens at 0.6GB. But the headline 2.9GB holds only on machines with spare OS page cache, an 8GB M1 MacBook hit 0.66, 0.81 tok/sec (30x slower) and still matched the memory claim. The model card admits agent capability is not yet ready. This video decodes the released checkpoint's exact file layout, the routing configuration that halves expert traffic, what quantization + LoRA actually cost (3.9 points on benchmarks), and why the speedup claim silently assumes your Mac has 20GB of free RAM sitting idle. For builders evaluating local inference on Apple Silicon, and anyone curious where the next generation of model efficiency is actually heading. Chapters: 0:00 The memory trick with an asterisk 1:13 Where the gigabytes actually went 2:12 One point three stays, everything else faults 3:34 Two hundred seventy MB per word 4:39 Why they halved the expert count 5:55 How prediction heads beat the disk 7:01 The adapters that salvaged quality 8:00 The benchmark nobody's spare RAM bought 9:15 Thirty times slower on real hardware 10:31 Why context didn't break it 11:53 Agent work: not ready yet 12:55 Running on phones, with warnings 14:05 When kernels lie to your chip 15:19 Does streaming actually work 16:23 The framework that changes the game Tools & resources mentioned: - Edge0 GitHub Repository: https://github.com/Edge0-AI/Edge0 - Edge0-35B Model Card: https://huggingface.co/Edge0/Edge0-35B-A3B-preview - Edge0 Documentation, Models: https://github.com/Edge0-AI/Edge0/blob/main/docs/models/edge0-35b.md - Edge0 Documentation, Streaming: https://github.com/Edge0-AI/Edge0/blob/main/docs/streaming.md - Edge0 Documentation, Prerouter: https://github.com/Edge0-AI/Edge0/blob/main/docs/prerouter.md About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #mixtureofexperts #localai #applesilic