1M Context In 500MB | DeepSeek V4.1 + TurboQuant

From the creator

Quantization cuts a million-token KV cache to 890MB, how DeepSeek V4.1 and TurboQuant compress attention memory, and what llama.cpp and vLLM let you use today. DeepSeek V4.1-Flash holds a million-token KV cache in 890 megabytes, roughly 437 times smaller than DeepSeek-V1, by combining compressed latent attention (storing fewer numbers per token across shared layers) with 4-bit FP4 quantization trained into the model itself. Google's TurboQuant compresses any model's cache post-training via random rotation, optimal scalar quantization, and 1-bit error correction, delivering no measurable loss at 3.5 bits and slight loss at 2.5 bits on 16-bit baselines. Stacking TurboQuant on V4.1's already-compressed cache trims 890MB to roughly 690MB (3.5-bit) or 500MB (2.5-bit) on paper, but no published implementation combines the two, and vLLM's TurboQuant support remains incomplete for DeepSeek's attention style. For models you run locally (Qwen3-8B, Llama 3.1, Ministral), the real leverage is 8-bit or 4-bit KV cache quantization. llama.cpp has run Hadamard rotation on quantized caches since April 2026; a single A100 benchmark showed -ctk q8_0 -ctv q8_0 saves 2.11GB at 32K context with 0.1% perplexity drift, though decode speed halves past 64K tokens. vLLM's TurboQuant modes deliver 2.3, 3.7× capacity but cost 40, 52% throughput; Red Hat's study recommends plain 8-bit as the default for reliable serving. This video breaks down where the 890-byte figure comes from, what TurboQuant adds, the gaps between paper math and working code, and which cache settings actually help in llama.cpp and vLLM. For builders optimizing local inference and API cost. Chapters: 0:00 Intro 1:10 DeepSeek 3:10 TurboQuant 6:14 llama.cpp 8:54 vLLM 10:12 Conclusion Tools & resources mentioned: - DeepSeek-V4.1-Flash: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash - llama.cpp: https://github.com/ggml-org/llama.cpp - vLLM: https://vllm.ai - DeepSeek API: https://api-docs.deepseek.com - TurboQuant (arXiv): https://arxiv.org/abs/2504.19874 - DeepSeek-V4.1-Flash Technical Report (arXiv): https://arxiv.org/abs/2609.19969 About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #quantization #kvcache #llmoptimization

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.