Bonsai 2: Qwen 27B on 6GB VRAM
PrismML says Bonsai 2 keeps 98% of Qwen3.8 27B's performance in a 6 GB file, so I tested it against the full model in the same agent harness, with thinking off on both. On short tasks it's a tie (48.6 vs 47.1), but on six long agentic builds the full Qwen scored 35 out of 60 while Bonsai didn't produce a single working app, at one point running the same search 114 times. Here's how ternary compression works, what their own whitepaper says about agentic benchmarks, and what to look for in the traces before you trust a compressed model with agent work. RESOURCES: Blogpost: https://prismml.com/news/bonsai-2-27b My voice to text App: whryte.com Website: https://engineerprompt.ai/ RAG Beyond Basics Course: https://prompt-s-site.thinkific.com/courses/rag Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off). 00:00 Bonsai2 01:02 How Bonsai2 Compresses 04:54 Benchmarks vs Agent Loops 05:31 Test Setup and Scoring 07:06 Short Task Results 09:10 Agentic Builds Breakdown 12:11 Why It Fails and Takeaways