Can One RTX 3090 End Your Monthly AI Bill?

From the creator

Qwen 3.8-27B on one RTX 3090: testing the GitHub repo claiming 1,000 tok/s at 64 concurrent users with vLLM, against Red Hat's 19x batching multiplier and the real per-user throughput. A GitHub repo (syv-ai/qwen38-27b-rtx3090) claims one RTX 3090 running Qwen 3.8-27B can deliver roughly 1,000 tokens per second serving 64 concurrent requests via vLLM, using int8 tensor-core GEMMs, fp16 DeltaNet state, MTP drafts and split-KV attention, but the headline splits into three separate claims: whether those numbers reproduce outside the author's environment, whether the advertised 262k-token context survives 64 concurrent KV-cache slots, and what per-user latency actually looks like under load. Red Hat's independent 2026 benchmark measured a 19x gap between vLLM (793 tok/s) and Ollama (41 tok/s) on the same hardware, collapsing to 20% difference at single-user load, isolating batching architecture as the real multiplier, not model-specific tuning. That's more than syv-ai claims, which itches. The repo's own patches may be doing far less work than the readme implies; no ablation has tested each technique separately, and no third-party rerun of the 64-concurrent figure exists yet. Once you split 1,000 tok/s across 64 seats, each user gets roughly 15 tokens per second, competitive with 2024 reports of Llama 3.1 8B on the same card but with a 27B model that needs heavy quantization to fit. The bigger lever turns out to be Qwen's xhigh reasoning preset, which burns 7-9x more thinking tokens than the low/medium presets while scoring nearly identically on benchmarks. For builders asking whether a used 3090 ends the monthly bill: the card handles the concurrency, but the real constraint is the thinking budget you can now control yourself instead of renting. True for anyone evaluating local Qwen 3.8 deployment on consumer GPUs. Chapters: 0:00 $2k GPU serving 64 people at once 0:35 Why patches shipped matters more than benchmarks 2:16 Single user leaves 90% of the card idle 3:47 Batching flips the entire economics 5:26 The 19x gap that labs actually measured 7:04 Two years of proof nobody talks about 8:44 27 billion squeezed into 24 gigabytes 10:13 Context windows don't survive shared seats 11:54 How 1000 tok/s becomes 15 per person 13:17 Same card, same model, three different speeds 14:43 The reproduction gap that should worry you 16:16 Thinking tokens burn 9x by default 18:05 The real lever isn't the hardware Tools & resources mentioned: - syv-ai/qwen38-27b-rtx3090: https://github.com/syv-ai/qwen38-27b-rtx3090 - vLLM: https://github.com/lm-sys/vllm - Ollama: https://ollama.ai - LM Studio - llama.cpp: https://github.com/ggerganov/llama.cpp - Qwen 3.8-27B: https://github.com/QwenLM/Qwen - GigaGPU vLLM setup guide About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #qwen 3.8 27b #vllm #ollama #homelab #llm inference

Choose to Build with AI
Matched to AI Engineering

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.