You Don't Need A New GPU For Qwen 3.8 27B

From the creator

Qwen 3.8 27B doubles speed with one llama.cpp flag, multi-token prediction was already inside your GGUF, just switched off by default. Qwen 3.8 27B ships with built-in multi-token prediction (MTP) layers baked into the model weights, but standard inference engines leave them disabled by default. Flipping two llama.cpp flags, --spec-type draft-mtp and --spec-draft-n-max 2, activates speculative decoding: a small head drafts tokens ahead while the main model verifies them in a single memory pass, cutting decode time roughly in half across most consumer GPUs without retraining or file changes. The community repo sudoingX/qwen38-mtp pooled benchmarks across 53+ cards spanning a decade of silicon, RTX 3090, 5090, older Pascal, Apple Silicon, and others, showing speedups ranging from +33% to +145% depending on your hardware, with acceptance rates and latency trade-offs varying by VRAM tier and concurrency. Unsloth's official Qwen 3.8 27B GGUF quantizations ship with MTP already enabled; their Q4_K_XL variant (18GB) targets 24GB VRAM cards with the feature active out of the box. On NVIDIA DGX Spark clusters, MTP alone lifted FP8 throughput from 7.9 to 17.1 tok/s (~2.16x), and when stacked with precision tuning and parallelism across 11 optimization rungs, the full stack achieved 4.5→77.3 tok/s. However, baseline measurements drift easily (concurrent request slots, browser memory pressure, sampler settings all shift the before number), and acceptance rate is a vanity metric, only paired A/B wall-clock speed matters. This video walks you through the flag, real speedup ranges by card, why Apple M4 saw no gain, the hidden cost to time-to-first-token, and how to verify MTP is actually running on your machine. For local LLM builders running Qwen on consumer or workstation GPU. Chapters: 0:00 The hidden layers already in your file 1:58 Why Qwen left its own speedup off 2:54 Guessing tokens, zero quality loss 4:20 How bandwidth becomes your hard limit 6:57 Real speeds across 53 actual cards 8:32 VRAM size changes the sweet spot 10:14 Why Apple M4 gained nothing 11:12 Acceptance rate is a trap 12:35 Unsloth made it easier than you think 13:58 One browser window halves your benchmark 15:09 Baseline measurement is fragile 16:16 When the engine moved faster than the flag 17:22 Recommended settings wipe the speedup 18:52 Tiny external model crushes the built-in head 20:29 The free rung on the optimization ladder 22:28 Should you actually type it Tools & resources mentioned: - llama.cpp: https://github.com/ggerganov/llama.cpp - Unsloth GGUF Qwen 3.8 27B: https://huggingface.co/unsloth - sudoingX/qwen38-mtp: https://github.com/sudoingX/qwen38-mtp - Ollama: https://ollama.ai - Open WebUI: https://github.com/open-webui/open-webui - Hugging Face - vLLM - SGLang About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #qwen38 #localllm #llmoptimization

Choose to Build with AI
Matched to Best Local LLM

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.