DeepSeek Shrunk Its Memory 437X
DeepSeek V4.1-Flash slashes KV cache to 890 bytes per token, 437x smaller than V1. See how mixture of experts, encoder-decoder split, and FP4 quantization cut attention memory for million-token contexts. DeepSeek's V4.1-Flash achieves a 437× reduction in attention memory footprint by redesigning what state a long-context model stores, not just shrinking formats. The 552B multimodal mixture of experts model supports up to one million tokens while keeping its global KV cache to roughly 890 bytes per token, achieved through a 40-layer causal encoder-decoder architecture (20 encoder, 20 decoder layers), Compressed Sparse Attention 2 (CSA2) for cross-layer index reuse, sliding-window attention bounded replay for local cache reconstruction, and FP4 (E2M1) KV quantization with per-16-channel scaling. Previously, DeepSeek-V2's Multi-Head Latent Attention cut KV memory by 93% using low-rank factorization, and V3.2-Exp added sparse attention to shift cost from quadratic to linear in selected tokens. This new approach stacks those gains further by splitting the model's read and answer paths, trading persistent memory for compute via on-demand replay. DeepSeek reports the decoder's persistent cache footprint drops to roughly one-eighth of V4-Flash's size on host memory or SSD. The model was pretrained on 45 trillion tokens and reaches production via Baseten's million-token API and vLLM's CUDA and Ascend deployment guides (Atlas 800 A2/A3). Official API pricing: $0.15 per million uncached input tokens, $0.60 per million output, $0.003 per million cached inputs (double at peak). However, these headline memory figures remain DeepSeek's self-reported claims; no independent third-party measurement of the 890-byte footprint has been published. NVIDIA's parallel NVFP4 KV cache technique reports under 1% accuracy loss on Blackwell, validating that FP4 quantization is production-viable, but does not verify DeepSeek's specific numbers. For builders evaluating long-context inference costs and AI researchers tracking attention optimization techniques. Chapters: 0:00 The 890-byte breakthrough 1:05 Why your cache explodes 2:53 How V2 already cut 93% 4:38 One model, two halves 6:32 Reusing indices, replaying locally 8:46 Four bits to hold it all 10:49 The math behind the bill 12:36 What actually works in production Tools & resources mentioned: - DeepSeek V4.1-Flash: https://www.modelscope.cn/models/deepseek-ai/DeepSeek-V4.1-Flash - vLLM Ascend Deployment: https://docs.vllm.ai/projects/ascend/en/latest/tutorials/models/DeepSeek-V4.1-Flash.html - Baseten: https://www.baseten.co - NVIDIA NVFP4 KV Cache: https://developer.nvidia.com/blog/optimizing-inference-for-long-context-and-large-batch-sizes-with-nvfp4-kv-cache About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #deepseek #kv cache #mixture of experts