The 80B Model That Fits In 4.3 GB Of RAM

From the creator

Mixture-of-experts architecture lets 80B parameter Qwen models run on Mac in 4.3 GB RAM via Swiftlet's disk-streaming runtime, and a 35B sibling on iPhone, though with caveats. Swiftlet is a Swift+Metal runtime that streams mixture-of-experts weights from SSD to run Qwen MoE models locally on Apple Silicon with minimal RAM. The 80B-parameter Qwen3-Next-80B-A3B activates only ~3B parameters per token across 512 routed experts, letting Swiftlet keep just the 2.5 GB dense core resident while fetching unused experts from disk on demand, a trick that trades speed for memory. On a base M5 MacBook (24GB unified memory), the full 80B model peaks at 4.3 GB RAM and decodes at 4.5, 5 tokens/second; on an iPhone 17, the 35B sibling (Qwen3.6-35B-A3B) runs at ~2.5 GB RAM and 1 tok/s after a 30-second first reply. The 35B ships in Priv AI via an 18 GB resumable download; the 80B requires 42 GB disk on Mac. Outside testing (a 2020 M1 Mac mini, a 64GB MacBook Pro) confirmed the cache doesn't bottleneck speed, the GPU dispatch is the constraint, and validated that a 397B model can stream at 1.4 tok/s on hardware that cannot hold it resident. The cost: storage (18, 42 GB), latency on long prompts (minutes for 591-token system prompt), and 3B-parameter recall ceiling despite 80B model fluency. For builders and AI-curious engineers evaluating local inference on Apple Silicon versus cloud APIs, this shows the RAM-speed tradeoff and when storage-streaming actually wins. Chapters: 0:00 80 billion parameters, 2.5 gigabytes of RAM 0:23 The M5 MacBook goes toe-to-toe with a server 2:01 Why only ten experts wake up per word 3:26 The hidden cost of loading every expert 5:02 One SSD read per word, forty-two gigabytes total 6:38 Why the disk isn't the bottleneck here 8:03 The iPhone gets the smaller sibling instead 9:33 Eighteen gigabytes lands on your phone today 10:58 A 2020 Mac mini breaks the maintainer's numbers 12:38 Big model talk, small model memory 14:13 When a laptop can't hold the model at all 16:00 A 397 billion parameter model on an iPhone Pro 17:39 What you actually get, machine by machine Tools & resources mentioned: - Swiftlet: https://github.com/leonickson1/Swiftlet - Priv AI: https://privy.app - Qwen3-Next-80B-A3B: https://huggingface.co/Qwen/Qwen3-Next-80B-A3B - MLX: https://github.com/ml-explore/mlx - ANEMLL About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #moe #apple #llm

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.