Can One 12GB GPU Run Qwen's 27B?

From the creator

Qwen3.8-27B fits a 12GB card at 11.8GB, but the math breaks it. Learn which quantized GGUF actually works in llama.cpp, Ollama, and LM Studio. Qwen3.8-27B is Alibaba's 27-billion-parameter vision-language model now available as a quantized GGUF at 11.8GB, but downloading the recommended build without understanding the real memory cost will leave you with just 24MB of usable cache. ISTA-DASLab, the team behind GPTQ, released eight model files using GSQ scalar quantization and RCO, an exact budget-constraint solver that assigns one of eleven different storage formats to each of 851 tensors, from full 16-bit down to 1.75-bit IQ1_M. The separate 931MB vision projector (mmproj) eats almost all your headroom, and on a 12GB card the recommended 3.50 bpw build leaves only 386 tokens of context at best. The 10.09GB three-bit variant instead scores 100.0 on AIME25 (matching the original), costs just half a point on GPQA Diamond, and leaves 26,000 tokens of headroom with images loaded. The video walks through exact byte math, why the architecture's 48 recurrent layers make a 12GB card viable at all, where the infrastructure bugs hide in llama.cpp and Ollama, and why RCO's allocation strategy isn't monotone in budget. Built for developers and AI builders deciding whether Qwen3.8 runs locally, what quantization buys you, and which of the eight files actually fits. Chapters: 0:00 Eight files hide behind the headline 1:11 Binary gigabytes versus decimal numbers 2:10 The vision projector wasn't compressed 3:15 Why DASLab chose the weaker method 4:34 How RCO spends your memory exactly 5:38 Eleven formats packed into 851 tensors 6:43 Half your file is just vocabulary 7:40 Two quantizations beat the original 8:54 Why 48 layers aren't attention layers 9:44 Every token costs 64 kilobytes 10:58 llama.cpp fails silently at 96K tokens 12:08 All three platforms crash or hang 13:28 The one format llama.cpp can't compute 14:28 24 megabytes left after loading 16:08 Download the 10GB version instead 17:07 The code they didn't ship yet Tools & resources mentioned: - ISTA-DASLab Qwen3.8-27B-GSQ-RCO-GGUF: https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF - llama.cpp: https://github.com/ggml-org/llama.cpp - Ollama: https://ollama.ai - LM Studio: https://lmstudio.ai - GSQ: Accurate Low-Bit Quantization (arXiv 2604.18556): https://arxiv.org/abs/2604.18556 - RCO: Model Compression with Exact Budget Constraints (arXiv 2605.00649): https://arxiv.org/abs/2605.00649 About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #qwen27b #localllm #quantization

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.