Which Qwen 27B Quant Should You Run?
On a 24 GB card, start with Q4_K_M, not Q6_K. A file that fits on your drive is not a coding session that fits in memory; here is what Q6 actually buys and how to test both on your own code. In Bartowski's Qwen3.8-27B GGUF release, the Q4_K_M file is 17.4 GB and the Q6_K file is 23.9 GB. Those are download sizes, not the video memory a run uses: a coding session also has to hold its context (project instructions, files, tool results and the cached attention that grows with the conversation). Q4 leaves far more room on a 24 GB card. Q6 is not impossible, but its margins are much narrower, and a model partly placed in system memory is a different way of running it. LM Studio's lms load --estimate-only gives an early estimate that follows your context length and GPU offload settings; the definitive check is watching real memory use under a representative coding workload. Quality: in Bartowski's own measurement (KL divergence against the bf16 model on wiki.test.raw, 100 chunks), Q6_K stays much closer to the original model's next-word predictions than Q4_K_M (mean KLD 0.0030 against 0.0139). That is a real fidelity difference on Wikipedia text, not proof of better code. Tool calls: neither quant was tested on matched tool calls here. Whether a tool call works depends on your host app as much as on the model, so check that the app parsed the request, ran it and returned the real file contents before judging the model. A tool call can run perfectly while the code it leads to is still wrong. No matched Q4 versus Q6 coding test exists here, so the video lays out one you can run on your own machine: a focused bug fix and a multi-file refactor, checks written in advance and kept outside the model's workspace, identical settings for both files, and five outcomes logged per run (tests passed, tools ran, time to a verified result, peak memory, manual fixes), repeated over several attempts. Verdict: on a 24 GB card, start with Q4_K_M. Step up to Q6_K only if your context fits comfortably and your own checked tasks improve enough to justify the extra memory and any measured time cost. Media credits: llama.cpp logo: github.com/ggml-org/llama.brand, CC BY-NC 4.0 (used with ggml-org's permission) Wikipedia logo: Wikimedia Foundation, CC BY-SA 3.0 (https://creativecommons.org/licenses/by-sa/3.0/) SSD photo: PantheraLeo1359531, Wikimedia Commons, CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) Stock footage (Pexels): RDNE Stock project, Efrem Efre Chapters: 0:00 Intro 0:57 Memory 2:52 Quality 4:04 Tool calls 5:32 Your test 9:10 Conclusion Tools & resources mentioned: - Qwen3.8-27B GGUF quants by bartowski (Hugging Face): https://huggingface.co/bartowski/Qwen3.8-27B-GGUF - Qwen3.8-27B official model page: https://huggingface.co/Qwen/Qwen3.8-27B - LM Studio docs: lms load (--estimate-only): https://lmstudio.ai/docs/cli/local-models/load - LM Studio docs: tool use: https://lmstudio.ai/docs/developer/openai-compat/tools - llama.cpp docs: function calling: https://github.com/ggml-org/llama.cpp/blob/master/docs/function-calling.md About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #qwen27b #quantization #localllm