Xiaomi & DeepSeek Solved A Massive Problem In AI... Allegedly
DeepSeek V4.1 Flash and Xiaomi HySparse2 cut KV-cache memory roughly 4, 4.5x, but million-token recall remains unverified, here's what the savings actually prove. DeepSeek and Xiaomi both published architectural fixes in September 2026 for the main bottleneck in long-context AI: the KV cache, the per-layer notes every model keeps on every token it reads. DeepSeek V4.1 Flash, live now on its API and open-weights on Hugging Face, compresses those notes to 890 bytes per token (roughly a quarter of V4 Flash) through Causal Encoder-Decoder architecture, Compressed Sparse Attention 2 (CSA2) that lets layers share notes, and FP4 quantization. At cache hit, off-peak pricing hits $0.003 per million tokens. Xiaomi's HySparse2 paper (arXiv 2609.26368, by Luo Fuli's MiMo team) calculates similar savings for its test model, 4.5x less KV cache and 5x less prefill compute at a million tokens through KV Bridging and KV Reuse, clean arithmetic rather than a measured run. The real question: whether either model can actually find information buried a million tokens back. Xiaomi's RULER-v2 tests stop at 256k tokens, where HySparse2 scored 58.45 out of 100. DeepSeek's report skips a million-token retrieval score entirely; independent testing via Artificial Analysis reaches only 100k tokens (84% on AA-LCR). DeepSeek's earlier V4 model showed degradation beyond 128k. Cost savings are proven math, but million-token recall remains unverified. For builders using Claude Code agents, AI engineers, and anyone evaluating million-token windows in production. Chapters: 0:00 Intro 1:16 The trick 4:17 Xiaomi 6:36 DeepSeek 9:34 Verdict Tools & resources mentioned: - DeepSeek V4.1 Flash: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash - DeepSeek V4.1 Flash Technical Report: https://arxiv.org/abs/2609.19969 - Xiaomi HySparse2 Paper: https://arxiv.org/abs/2609.26368 - DeepSeek API: https://api-docs.deepseek.com/quick_start/pricing - Artificial Analysis Long-Context Reasoning About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #deepseekv4 #ai #kvcache #longcontext #transformers