DeepSeek Shrunk Its Memory 437X

From the creator

DeepSeek V4.1-Flash slashes KV cache to 890 bytes per token, 437x smaller than V1. See how mixture of experts, encoder-decoder split, and FP4 quantization cut attention memory for million-token contexts. DeepSeek's V4.1-Flash achieves a 437× reduction in attention memory footprint by redesigning what state a long-context model stores, not just shrinking formats. The 552B multimodal mixture of experts model supports up to one million tokens while keeping its global KV cache to roughly 890 bytes per token, achieved through a 40-layer causal encoder-decoder architecture (20 encoder, 20 decoder layers), Compressed Sparse Attention 2 (CSA2) for cross-layer index reuse, sliding-window attention bounded replay for local cache reconstruction, and FP4 (E2M1) KV quantization with per-16-channel scaling. Previously, DeepSeek-V2's Multi-Head Latent Attention cut KV memory by 93% using low-rank factorization, and V3.2-Exp added sparse attention to shift cost from quadratic to linear in selected tokens. This new approach stacks those gains further by splitting the model's read and answer paths, trading persistent memory for compute via on-demand replay. DeepSeek reports the decoder's persistent cache footprint drops to roughly one-eighth of V4-Flash's size on host memory or SSD. The model was pretrained on 45 trillion tokens and reaches production via Baseten's million-token API and vLLM's CUDA and Ascend deployment guides (Atlas 800 A2/A3). Official API pricing: $0.15 per million uncached input tokens, $0.60 per million output, $0.003 per million cached inputs (double at peak). However, these headline memory figures remain DeepSeek's self-reported claims; no independent third-party measurement of the 890-byte footprint has been published. NVIDIA's parallel NVFP4 KV cache technique reports under 1% accuracy loss on Blackwell, validating that FP4 quantization is production-viable, but does not verify DeepSeek's specific numbers. For builders evaluating long-context inference costs and AI researchers tracking attention optimization techniques. Chapters: 0:00 The 890-byte breakthrough 1:05 Why your cache explodes 2:53 How V2 already cut 93% 4:38 One model, two halves 6:32 Reusing indices, replaying locally 8:46 Four bits to hold it all 10:49 The math behind the bill 12:36 What actually works in production Tools & resources mentioned: - DeepSeek V4.1-Flash: https://www.modelscope.cn/models/deepseek-ai/DeepSeek-V4.1-Flash - vLLM Ascend Deployment: https://docs.vllm.ai/projects/ascend/en/latest/tutorials/models/DeepSeek-V4.1-Flash.html - Baseten: https://www.baseten.co - NVIDIA NVFP4 KV Cache: https://developer.nvidia.com/blog/optimizing-inference-for-long-context-and-large-batch-sizes-with-nvfp4-kv-cache About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #deepseek #kv cache #mixture of experts

Choose to Build with AI
Matched to AI Explained

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.