Do You Need A $10,000 PC To Run A Huge Qwen Model?

From the creator

Qwen 3.8 27B runs on a single RTX 3090 because only 6B parameters activate per word. Here's why a $2K setup beats the $10K myth. The Qwen3.8-Flash-Next model stores roughly 177 billion parameters but activates only 6 billion per token, the rest sit idle until needed. NVIDIA's architecture splits this into 125B total (6B activated), 51B n-gram embedding lookup table, and 4B multi-token prediction, with 512 experts routing just 11 active per word. That lookup table was designed by Qwen to be offloadable, making this massive model surprisingly efficient on consumer hardware. The real bottleneck isn't VRAM capacity, it's memory bandwidth. An RTX 3090 reads at 900GB/s, enabling token generation even when weights must stream from system RAM. llama.cpp splits the model automatically between your GPU and DDR4, requiring roughly 140GB combined (not all VRAM). A used RTX 3090 costs ~$1,032; a 128GB DDR4 setup runs ~$960; total build lands around $2,000. The catch: system memory prices spiked through 2026. Running baseline inference at full quality on one consumer card means accepting 16, 28 tokens/second; software optimizations like speculative decoding push it to 66 tokens/second but diverge outputs. This video walks through parameter activation, bandwidth vs. capacity, real benchmark splits (prompt-processing vs. token-generation), and the actual cost math, for builders weighing local inference against cloud costs. Chapters: 0:00 Why 177B parameters need only 6B to work 1:30 The 51 billion parameter lookup trick 2:30 Waking only 11 experts per word 3:20 How bandwidth beats raw compute power 4:49 Why the 3090 still dominates at reading 5:53 One card, four times slower over time 7:01 Trading speed for memory and context 8:31 Four cards barely hold it anyway 10:02 The one card that actually fits it 11:42 Speeding up 60% with one flag 13:05 Why RAM matters more than VRAM 14:20 DDR4 prices broke the budget math 15:34 Build for total memory, not the card Tools & resources mentioned: - llama.cpp: https://github.com/ggerganov/llama.cpp - Qwen3.8-Flash-Next: https://huggingface.co/Qwen/Qwen3.8-Flash-Next - NVIDIA Qwen3.8-Flash-Next-NVFP4: https://huggingface.co/nvidia/Qwen3.8-Flash-Next-NVFP4 About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #qwen #local llm #ai hardware

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.