How To Run A 27B AI On A 16GB GPU At 100K Tokens/S [NInfer]

From the creator

Run Qwen 3.8 27B with a 100K-token context on one 16 GB RTX 4080: NInfer's 3-bit GSQ model, compressed conversation memory and custom CUDA engine explained. Yes, a 27-billion-parameter model with 100,000-token context fits on a single 16GB RTX 4080. Developer roofkid published NInfer-4080, a specialized CUDA inference engine running Qwen 3.8 27B compressed to 3 bits using GSQ quantization from ISTA-DASLab. The reported speeds: about 2,720 tokens/second reading short prompts and about 1,970 at the full 100K (under a minute for a 100K prompt), and about 150 to 200 tokens/second writing code (around 100 on prose) in daily use. The 262 tokens/second headline came from benchmark text that repeats itself; a real coding prompt measured 136. On the same card it reads nearly twice as fast as llama.cpp, and the developer reports it ahead on writing at every length, by a smaller margin. This breakdown covers how 27 billion weights plus 100K tokens actually fit into 16GB (model compression to 3 bits, KV cache compression to ~2GB), what those speed claims really mean (prefill vs. generation, speculative decoding, benchmark text), and what it costs: the engine rejects other GPU architectures at compile time, requires 64-bit Linux + CUDA 13.1+, runs one request at a time, and reserves 5.5GB of pinned system memory. One developer's transparent measurements on one card, with no independent reproduction yet, but honest about where test numbers don't match real use. For builders working with massive documents on an RTX 4080, and anyone curious how modern coding agents (this one was built by directing DeepSeek V4.1 Flash inside the Pi agent tool, for about $13 of API spend) can beat general-purpose inference software. Chapters: 0:00 Intro 1:27 Memory 3:40 Speed 6:15 Comparison 8:43 Builder 10:01 Your card 12:13 Setup 13:59 Conclusion Tools & resources mentioned: - NInfer-4080: https://github.com/roofkid/ninfer-4080 - Qwen 3.8 27B GSQ (ISTA-DASLab): https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ - Qwen 3.8 27B NInfer (roofkid): https://huggingface.co/roofkid/Qwen3.8-27B-GSQ3-NInfer - llama.cpp: https://github.com/ggml-org/llama.cpp - NInfer-4090 (upstream fork): https://github.com/sergiuszm/ninfer-4090 - NInfer-5060 Ti (community fork): https://github.com/freedomNTD/ninfer-5060ti About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #quantization #local llm #qwen

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.