You Can RunsA 180B Qwen AI Model On A 12GB GPU

From the creator

Run Alibaba's 180B Qwen3.8-Flash-Next on a 12GB GPU using quantization, MoE routing, and smart memory tiering across VRAM, RAM, and SSD storage. Alibaba's Qwen3.8-Flash-Next is a 180-billion-parameter mixture-of-experts model, and a community engine called Strata runs it on a desktop with a 12GB graphics card by leaning on system RAM and an SSD. Strata is an open-source community project (MIT License), not an official Alibaba release. The graphics card holds the always-used base layers and a cache of the most frequently called experts; all the routed experts sit in system RAM, where the CPU computes the ones that aren't cached; and the 28.8GB lookup table stays on the SSD, which only reads the few rows each token needs. The model has a 125B backbone with about 6B parameters active per token, plus a 51B lookup table and a 4B drafting head. On one documented PC (RTX 5070 12GB, Ryzen 5 7600, 64GB of RAM), the Strata developer reports 90.3 tokens per second at a 4K context and 67.2 at 128K with the Q2_0 package (engine 0.1.14, run of September 28). Those are the developer's own numbers from one machine, not an independently audited benchmark. The quality scores come from a separate evaluation by ISTA-DASLab, the team that made the packages: on LiveCodeBench v6, Q2_0 scores 81.1% against the BF16 baseline's 87.4%, while the slightly larger IQ3_XXS scores 86.3% but generates 45.8 tokens per second at 128K in the developer's run. If you already own a matching desktop, it's a plausible weekend experiment, not a reason to buy new hardware or drop a dependable cloud workflow: run both packages on a coding problem with known tests and measure the wait for the first token, the time to a correct fix, and everyday stability. Chapters: 0:00 Intro 0:53 The model 2:33 Memory 4:10 Speed 6:13 Quality 7:48 Your PC 8:48 Conclusion Tools & resources mentioned: - Strata (community inference engine, MIT License): https://github.com/Niko1221/Strata - The Strata developer's speed run (engine 0.1.14, September 28): https://github.com/Niko1221/Strata/blob/main/bench/results/2026-09-28-speed-0114/README.md - Qwen3.8-Flash-Next model card: https://huggingface.co/Qwen/Qwen3.8-Flash-Next - ISTA-DASLab's GSQ-RCO GGUF packages (Q2_0, IQ3_XXS) and LiveCodeBench v6 results: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #quantization #mixture of experts #local LLM

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.