Can One RTX 3090 Run A Frontier Model?

From the creator

Run DeepSeek V4 Flash on a single RTX 3090 with FreeToken's MoE engine, 95% of the frontier model sleeps in RAM while 10.5 tokens per second flow live. DeepSeek V4 Flash is a 284-billion-parameter frontier model that works on a single 24GB graphics card because 95% of it is optional: only 8.28GB of the 166.88GB checkpoint is always-on, and the rest, 158.60GB of expert weights, lives in system RAM. FreeToken, an open-source mixture-of-experts serving engine from the Berkeley/MIT team behind vLLM and SGLang, treats your gaming PC as a unified inference platform by activating only the experts needed for each token. On a real RTX 3090, you hit 10.5 tokens per second in Open WebUI, but the card itself is not your constraint. The floor is 156, 168GB of system RAM (not 128GB), the prefill cost is a 5-second transfer of 140GB across PCIe 4.0 per long prompt, and the PCIe link itself runs at 78% utilization while the CPU sleeps. The paper from Keutzer, Zaharia, Stoica (Berkeley) and Song Han (MIT) compares FreeToken against llama.cpp, Ollama, and KTransformers on real agentic workloads, math, code issues, calendar agents, and finds a 1.3× throughput edge on the 3090 and up to 2.3× on newer hardware. It also ships an open defect: the calibration tool silently picks a backend 8.3× slower than the one it rejects. For builders asking whether a single consumer card can serve a real frontier model, the answer is yes, but only if you understand that the card is the least of your problems. Chapters: 0:00 The model that sleeps through generation 0:26 FreeToken's datacenter promise 1:32 Only 5% stays awake 2:30 10.5 tokens per second, on camera 3:24 When 128GB becomes the problem 4:31 No desktop tested here 5:45 Prefill wakes the entire model 6:40 140GB moves every long prompt 7:36 PCIe slots matter more than cards 8:45 The bottleneck nobody blames 10:15 Laptops hit 92% desktop speed 11:10 The real engines, fairly matched 12:27 Which stall destroys agent calls 13:48 Calibration picks the slower path 15:37 Same bug, newer hardware 17:07 Chat speed is the ceiling 18:06 The card that finally worked Tools & resources mentioned: - FreeToken: https://github.com/FlashML-org/FreeToken - DeepSeek V4 Flash (Hugging Face): https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 - FreeToken Research Paper (arXiv): https://arxiv.org/abs/2608.16157 - Open WebUI - llama.cpp - Ollama - KTransformers About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #FreeToken #MoE #DeepSeek

Choose to Build with AI
Matched to Continuous Batching

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.