Xiaomi & DeepSeek Solved A Massive Problem In AI... Allegedly

From the creator

DeepSeek V4.1 Flash and Xiaomi HySparse2 cut KV-cache memory roughly 4, 4.5x, but million-token recall remains unverified, here's what the savings actually prove. DeepSeek and Xiaomi both published architectural fixes in September 2026 for the main bottleneck in long-context AI: the KV cache, the per-layer notes every model keeps on every token it reads. DeepSeek V4.1 Flash, live now on its API and open-weights on Hugging Face, compresses those notes to 890 bytes per token (roughly a quarter of V4 Flash) through Causal Encoder-Decoder architecture, Compressed Sparse Attention 2 (CSA2) that lets layers share notes, and FP4 quantization. At cache hit, off-peak pricing hits $0.003 per million tokens. Xiaomi's HySparse2 paper (arXiv 2609.26368, by Luo Fuli's MiMo team) calculates similar savings for its test model, 4.5x less KV cache and 5x less prefill compute at a million tokens through KV Bridging and KV Reuse, clean arithmetic rather than a measured run. The real question: whether either model can actually find information buried a million tokens back. Xiaomi's RULER-v2 tests stop at 256k tokens, where HySparse2 scored 58.45 out of 100. DeepSeek's report skips a million-token retrieval score entirely; independent testing via Artificial Analysis reaches only 100k tokens (84% on AA-LCR). DeepSeek's earlier V4 model showed degradation beyond 128k. Cost savings are proven math, but million-token recall remains unverified. For builders using Claude Code agents, AI engineers, and anyone evaluating million-token windows in production. Chapters: 0:00 Intro 1:16 The trick 4:17 Xiaomi 6:36 DeepSeek 9:34 Verdict Tools & resources mentioned: - DeepSeek V4.1 Flash: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash - DeepSeek V4.1 Flash Technical Report: https://arxiv.org/abs/2609.19969 - Xiaomi HySparse2 Paper: https://arxiv.org/abs/2609.26368 - DeepSeek API: https://api-docs.deepseek.com/quick_start/pricing - Artificial Analysis Long-Context Reasoning About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #deepseekv4 #ai #kvcache #longcontext #transformers

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.