Bigger Graphics Card Or More RAM For A 125B Model

From the creator

Qwen 3.8 Flash Next needs 96GB RAM, not a graphics card, but which oversized weights actually belong where matters more than the total parameter count. Qwen 3.8 Flash Next claims to run on 96GB of system RAM with no GPU VRAM required, but the model architecture separates active computation from lookup tables and prediction heads in ways that flip the usual graphics-card-versus-memory trade-off on its head. The 125B headline parameter count masks 6B active per token, a separate 51B N-gram embedding table, and a 4B prediction head, roughly 180B stored, ~6B computed. bartowski's GGUF quantization hits 119.6GB on disk, larger than the stated memory requirement, because pieces are meant to offload to system RAM or SSD via mmap. DataCamp ran the full model in VRAM on a single RTX PRO 6000 wired into a coding agent; Phi-3.5-MoE and DeepSeek-MoE offload benchmarks show 99.9% VRAM savings paired with generation speed collapsing several times over and time-to-first-token stretching past 2 seconds. The hidden cost is cache hit rate: mixture-of-experts weights get guessed and multiplied, paying a penalty every token, while embedding tables get addressed and read with a known path that never guesses. The difference between an N-gram lookup and expert selection is the difference between a predictable SSD read and a cache miss on every word. LTT Labs measured premium DDR5 kits at $300, $500 over baseline; motherboard memory channels often matter more than raw gigabytes. This model is for builders deciding whether to buy a $5k+ workstation card or upgrade system memory, and for coding agents, where per-step latency turns a tolerable wait into an afternoon tax. Read the model card first: if documentation names a specific component and where it belongs, somebody designed for this. If it just talks about pushing layers off the card, you're inheriting every penalty from community offload benchmarks. For pure conversation, buy the memory. For agents, buy the card and stop pretending cheap was ever going to be fast. Chapters: 0:00 The 96GB Spec That Skips the Card 0:30 Where Six Billion Weights Actually Do Math 1:57 119GB on Disk vs 96GB in RAM 3:13 All-VRAM Runs on Enterprise Hardware 4:27 Why 730ms Beats Ten Seconds 5:42 The Card Barrier Nobody Mentions 7:01 The 99.9% VRAM Hack and Its Cost 8:17 When Moving Experts Tanks Your Speed 9:23 Cache Misses, Explained 10:32 Premium Memory Costs Real Money 11:41 Lookups Never Guess, Experts Always Do 13:28 The One-Sentence Model Card Test 14:36 Why Nobody Tested This Yet 15:57 The Per-Step Wait That Kills Agents 17:14 Buy RAM, Unless You Run Agents Tools & resources mentioned: - Qwen 3.8 Flash Next: https://huggingface.co/Qwen/Qwen3.8-Flash-Next - Unsloth Qwen 3.8 Next Docs: https://unsloth.ai/docs/models/qwen3.8-next - llama.cpp: https://github.com/ggerganov/llama.cpp - LM Studio: https://lmstudio.ai - Ollama: https://ollama.ai - Qwen Official Docs: https://qwen.readthedocs.io/en/latest/run_locally/llama.cpp.html - local-ai-zone Qwen 3.8 Deep Dive: https://local-ai-zone.github.io/blog/qwen3-8-flash-next-deep-dive.html About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #qwen3.8 #localai #mixtureofexperts

Choose to Build with AI
Matched to AI Explained

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.