You Can RunsA 180B Qwen AI Model On A 12GB GPU
Run Alibaba's 180B Qwen3.8-Flash-Next on a 12GB GPU using quantization, MoE routing, and smart memory tiering across VRAM, RAM, and SSD storage. Alibaba's Qwen3.8-Flash-Next is a 180-billion-parameter mixture-of-experts model, and a community engine called Strata runs it on a desktop with a 12GB graphics card by leaning on system RAM and an SSD. Strata is an open-source community project (MIT License), not an official Alibaba release. The graphics card holds the always-used base layers and a cache of the most frequently called experts; all the routed experts sit in system RAM, where the CPU computes the ones that aren't cached; and the 28.8GB lookup table stays on the SSD, which only reads the few rows each token needs. The model has a 125B backbone with about 6B parameters active per token, plus a 51B lookup table and a 4B drafting head. On one documented PC (RTX 5070 12GB, Ryzen 5 7600, 64GB of RAM), the Strata developer reports 90.3 tokens per second at a 4K context and 67.2 at 128K with the Q2_0 package (engine 0.1.14, run of September 28). Those are the developer's own numbers from one machine, not an independently audited benchmark. The quality scores come from a separate evaluation by ISTA-DASLab, the team that made the packages: on LiveCodeBench v6, Q2_0 scores 81.1% against the BF16 baseline's 87.4%, while the slightly larger IQ3_XXS scores 86.3% but generates 45.8 tokens per second at 128K in the developer's run. If you already own a matching desktop, it's a plausible weekend experiment, not a reason to buy new hardware or drop a dependable cloud workflow: run both packages on a coding problem with known tests and measure the wait for the first token, the time to a correct fix, and everyday stability. Chapters: 0:00 Intro 0:53 The model 2:33 Memory 4:10 Speed 6:13 Quality 7:48 Your PC 8:48 Conclusion Tools & resources mentioned: - Strata (community inference engine, MIT License): https://github.com/Niko1221/Strata - The Strata developer's speed run (engine 0.1.14, September 28): https://github.com/Niko1221/Strata/blob/main/bench/results/2026-09-28-speed-0114/README.md - Qwen3.8-Flash-Next model card: https://huggingface.co/Qwen/Qwen3.8-Flash-Next - ISTA-DASLab's GSQ-RCO GGUF packages (Q2_0, IQ3_XXS) and LiveCodeBench v6 results: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #quantization #mixture of experts #local LLM