A Tiny Graphics Card Runs A Giant AI Model
AirLLM runs 70B models on 4GB GPUs by streaming layers, but 292 seconds per token reveals the speed vs. memory tradeoff for local AI. AirLLM is an open-source Python library that solves a classic constraint: fitting a 70-billion-parameter model on a 4GB graphics card without quantization or pruning. Instead of loading the entire checkpoint into VRAM, it streams one decoder layer at a time from disk, keeping only ~1.75GB in memory per forward pass. The trick is real, but the GitHub repository's headline table omits the metric that matters: time. The maintainer, Gavin Li, measured a 2.8-trillion-parameter mixture-of-experts model (Kimi K3) on an RTX 6000 Ada and logged 3.72GB peak VRAM during generation, proof the mechanism works. He also logged 292 seconds per token, disk-bound. That's 4 minutes 52 seconds per word. Compression settings promised as a speedup actually slow the pipeline 8× by shifting the bottleneck from disk I/O to dequantization overhead. A smaller 3B model on an RTX 5060 Ti clocked 28.6 s/token, slower than a 32B run on the same hardware. The real constraint isn't VRAM or model size, it's the data pipeline: disk read speed (~210 MB/s measured), system RAM staging, and PCIe transfer overhead must repeat for every layer of every token. Turning off streaming entirely on models that fit yields 322× speedup, confirming the architectural tax. Progress logging is missing, so a frozen process and a running one look identical on screen. For builders running private data under compliance constraints or researchers experimenting offline, AirLLM unlocks models that otherwise would never fit. For interactive local-AI work, the speed penalty is prohibitive. This video is for anyone weighing local inference against cloud APIs and builders curious about what memory-streaming actually costs. Chapters: 0:00 Why 80 layers fit one at a time 1:14 Storage cost vs memory saved 2:46 Sparse models need less VRAM 4:07 The prefetch and release loop 5:07 Where the benchmarks actually are 7:01 The number nobody saw on the homepage 8:32 One hour costs 416 days of compute 9:52 Smaller models run slower somehow 11:11 Disk reads hit a hard ceiling 12:32 Compression makes it eight times worse 14:16 Loading normally: 322× faster 15:23 Who actually pays this penalty 16:27 Why progress looks like frozen 17:52 The table that should ship with it Tools & resources mentioned: - AirLLM: https://github.com/lyogavin/airllm - Llama 3.1: https://www.meta.com - Kimi K3: https://www.moonshot.cn - Qwen 2.5: https://github.com/QwenLM/Qwen About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #LocalAI #AirLLM #LLM