Qwen 3.8 Flash Next Can Run On Any Hardware Now... Almost

From the creator

Victoria cuts 44% of Qwen 3.8 Flash Next's routed experts. Here is what it really needs to run, what was removed, and how much coding ability it gives up. Victoria is a compressed version of Qwen 3.8 Flash Next by Farpoint (Hugging Face: rmonsurate). It removes 44% of the routed experts, 512 down to 288 per layer, while the router still picks 10 experts for every token, so the stored pool shrinks without a proportional speed gain. The publisher reports 34 to 38 tokens per second on a 128 GB M3 Max with the GGUF build, one request at a time with the draft head on (about 27 with it off). That is one tested machine, not a hardware minimum, and output speed says nothing about task time or whether the code works. Memory: the GGUF build keeps about 53 GB of weights resident while llama.cpp reads the model's lookup table from disk or system memory; the download is about 107 GB, or 155 GB with the full precision table, and on a Mac the table still uses the same physical memory pool. On one B300 at the same 4 bit precision and 8K context, the publisher measured process GPU memory falling from 82.2 GB for the unpruned parent to 49.9 GB for Victoria. The NVFP4 build for vLLM on Blackwell GPUs keeps both the 48.0 GiB of weights and the 95.4 GiB lookup table on the GPU, over 140 GiB before working memory. Coding: on Terminal-Bench 2.1 the first release GGUF solved 67 of 89 tasks against 79 for the full 16 bit parent in one matched run, the cost of the whole recipe (pruning, 4 bit quantization and retraining), not pruning alone. The 70.04% figure is a different checkpoint, the newer NVFP4 build averaged over three runs. A harsher 144 expert prune (Vernon) only recovered to 50.6% on HumanEval after 300 training steps. Setup: the GGUF files need the publisher's patched llama.cpp (b11276) for the draft head's 32 extra tensors; Maple, a Canadian fine tune of Victoria, ships NVFP4 only. Verdict: Victoria widens local access, but not to almost any hardware. If your setup already fits the full model and you need every solved task, keep the parent; NVFP4 needs Blackwell class hardware; if you already own high capacity hardware, test the GGUF build on your own tasks and tests before buying gear or replacing a working model. The Stack did not run the model; every result shown is the publisher's report. Sources and media credits: Model card, builds table, benchmarks and file listing: rmonsurate/Victoria and rmonsurate/Maple on Hugging Face (Farpoint) Patched llama.cpp builds: rmonsurate/llama.cpp release victoria-mtp-b11276 llama.cpp marks: github.com/ggml-org/llama.brand SSD photo: PantheraLeo1359531, Wikimedia Commons, CC BY 4.0 Product images: Apple (MacBook Pro with M3 Max), NVIDIA (DGX B300, GeForce RTX 5090), Kingston (DDR5) Stock footage (Pexels): Videas Cl, Dominiquemel16 Ramos Chapters: 0:00 Intro 0:52 Mac result 2:09 Memory 4:06 Smaller 6:16 Setup 8:07 Coding 10:04 Conclusion Tools & resources mentioned: - Victoria model card (Hugging Face): https://huggingface.co/rmonsurate/Victoria - Patched llama.cpp builds (b11276): https://github.com/rmonsurate/llama.cpp/releases/tag/victoria-mtp-b11276 - llama.cpp: https://github.com/ggml-org/llama.cpp - vLLM: https://github.com/vllm-project/vllm - Qwen 3.8 Flash Next (parent model): https://huggingface.co/Qwen/Qwen3.8-Flash-Next - Maple (Hugging Face): https://huggingface.co/rmonsurate/Maple - REAP (expert pruning method): https://github.com/CerebrasResearch/reap - Terminal-Bench: https://www.tbench.ai/ About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #LocalAI #Quantization #CodingAI

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.