Context Language Models: Your Agent Doesn't Need Compaction
Context Language Models (CLM) are a new approach to context management for AI agents from Meta Superintelligence Labs and the University of Washington. Instead of compacting or summarizing the conversation when the context window fills up, the agent edits its own context like a file: it shortens old tool outputs, replaces them with notes, and decides what is worth keeping. In this video I explain how Context Language Models work, why summaries like /compact lose useful information, and then test it on my own DGX Spark with Pi, the pi-clm extension, and Qwen3.8-27B running locally. I also cover the gotchas, including what self-editing does to your prefix cache. Let me know in the comments how you manage the context window in your own agents. Paper: https://arxiv.org/abs/2609.37725 Code: https://github.com/facebookresearch/context-language-models pi-clm (Pi extension): https://github.com/lolipopshock/pi-clm Pi coding agent: https://www.npmjs.com/package/@earendil-works/pi-coding-agent My voice to text App: whryte.com Website: https://engineerprompt.ai/ RAG Beyond Basics Course: https://prompt-s-site.thinkific.com/courses/rag Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off). 00:00 - Context Language Models 01:52 - Why Summaries and /compact Lose Context 03:44 - How Context Language Models Work 06:05 - Testing on a DGX Spark: Pi Summaries vs pi-clm 10:15 - Gotchas: Prefix Caching, Harness & Verdict