Claude Grade Coding From An 8.4GB Local AI Model

From the creator

Qwen 3.8 27B quantized to 8.4GB via GSQ-RCO still loses 9 coding points, here's what the benchmarks actually prove vs. Claude. Qwen3.8-27B is a 27-billion-parameter dense vision-language model from Alibaba's Qwen team that requires 53.8GB at full precision, but ISTA-DASLab's GSQ-RCO quantization technique reduces it to a single 8.4GB GGUF file by applying learned per-weight bit-depth selection and Riemannian constrained optimization across ~851 tensors. The lab claims 100.3% zero-shot recovery on multiple-choice benchmarks, but the script reveals that claim measures only quiz scoring, not real code generation. On LiveCodeBench, actual programming contest problems, the full model scores 85.7, the 11.8GB file matches it exactly, but the 8.4GB file drops to 76.6, losing about nine points on real coding tasks. Independent testing from ByteShape in llama.cpp found the 8.4GB file tied an ordinary Unsloth GGUF of the same size, contradicting the lab's claimed 4.6-point lead; the lab itself acknowledges ~90% of its compression tuning used English text, which explains real-world regressions in non-English prompts. The full Qwen model challenges Claude Opus only inside a multi-agent manager setup that burns 18.6M tokens per coding task, about 15× more than Claude's 1.2M, and none of these comparisons tested the 8.4GB file. For a 16GB GPU, download the 11.8GB IQ3_S file; for 12GB, choose between the 10.1GB file for better coding or the 8.4GB file for longer conversations. This deep dive is for builders evaluating local AI options and anyone curious whether quantization can actually preserve coding ability at extreme compression ratios. Chapters: 0:00 Intro 1:11 Basics 3:04 Lab scores 5:00 Outside tests 7:13 Versus Claude 10:40 Which file 14:01 Conclusion Tools & resources mentioned: - ISTA-DASLab Qwen3.8-27B GSQ-RCO GGUF: https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF - llama.cpp: https://github.com/ggerganov/llama.cpp - Ollama - LM Studio - LiveCodeBench - ByteShape - Unsloth - vLLM About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #qwen #quantization #local ai

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.