Don't Use Claude, Use This 320B Local Coding AI Instead

From the creator

IQuest-Q1 has 320 billion parameters, but only about 15 billion work on each token. Here is what that does and does not mean for self-hosting it as a coding agent. IQuestLab has released open weights for IQuest-Q1, a sparse mixture-of-experts model with 320 billion parameters in total. Each token is routed to 8 of its 256 experts, so roughly 15 billion parameters do the work per token (the lab's estimate). That cuts the math per token, but it does not make Q1 a 15B model: the whole model still has to be held, and its official Hugging Face repository lists about 650 GB of files. A hypothetical 4-bit version works out to about 160 GB of raw weights before the serving software and the conversation context; that is arithmetic, not a published file. The lab's own SGLang and vLLM examples split the model across eight GPUs (tensor parallel size 8), an example configuration rather than a tested minimum, and real response times depend heavily on your serving hardware and software. The video walks through two of the lab's reported repair cases (an extra whitespace separator that stopped earlier conversation turns from counting toward training, and a broken test environment whose crashes the grader blamed on the patches) and the lab's Terminal-Bench 2.1 table, where it reports 83.2 for Q1 against 89.1 for Claude Opus 5. The table mixes public scores with the lab's own runs, the agent benchmark tests the whole setup around the model, and the eight-hour limit per problem is a time budget, not a speed. The Stack did not run the model; every result shown is the lab's own report. Verdict: if you already operate multi-GPU servers and need local control over your code, Q1 is worth a bounded pilot: one bug with regression tests, then record correct tested code, wall-clock time and manual fixes. On a standard desktop, do not buy hardware for the 15B headline. Sources and media credits: IQuest Research: technical report, case recordings and benchmark chart (iquestlab.github.io) Model card, deployment examples and file listing: IQuestLab/IQuest-Q1 on Hugging Face SSD photo: PantheraLeo1359531, Wikimedia Commons, CC BY 4.0 GPU images: NVIDIA Stock footage (Pexels): Kelly Chapters: 0:00 Intro 0:55 Coding 2:38 15B active 4:10 Memory 6:32 Scores 7:54 Run it 9:55 Conclusion Tools & resources mentioned: - IQuest-Q1 model card (Hugging Face): https://huggingface.co/IQuestLab/IQuest-Q1 - IQuest-Q1 technical report: https://iquestlab.github.io/ - IQuest-Q1 on GitHub: https://github.com/IQuestLab/IQuest-Q1 - SGLang: https://github.com/sgl-project/sglang - vLLM: https://github.com/vllm-project/vllm - Claude Code: https://github.com/anthropics/claude-code - Codex CLI: https://github.com/openai/codex - Terminal-Bench: https://www.tbench.ai/ About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #localai #codingai #llmarchitecture

Choose to Build with AI
Matched to AI Agents

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.