The Local AI Coding Breakthrough We Need? (Bonsai 2)
Bonsai 2 27B squeezes Alibaba's Qwen 3.8 into a 6 GB file and matches it on quick one-prompt coding, but on agent work it keeps about three-quarters of the score. A free local breakthrough in size, not yet a paid-tool replacement. PrismML's Ternary Bonsai 2 27B (Apache 2.0, 5.93 GB) stores every weight as -1, 0 or +1, squeezing Alibaba's full 53.8 GB Qwen 3.8-27B model to about a ninth of its size while keeping 98.2% of its average score across PrismML's own 20-benchmark suite. For one-prompt coding jobs, write a function, explain a snippet, get one answer, that holds: on LiveCodeBench v6 (1,055 unit-tested problems) Bonsai 2 scores 90.07 against 90.05 for the full model, beating a conventional 2-bit quantization (70.05). But multi-step agent work, where the AI reads files, edits code, runs terminal commands and loops until done, as Claude Code and Cursor do, drops to three-quarters: on SWE-bench Verified (500 real GitHub issues) it resolves 60.8% vs the full model's 80.6%; on Terminal-Bench 2.1 it scores 52.8 vs 69.7. On SWE-bench that is roughly where Claude 3.7 Sonnet, the model behind Claude Code at its February 2025 launch, scored in Anthropic's own table (62.3%). With a long agent context it needs roughly 13 GiB of memory, so a 16 GB graphics card or a bigger Mac (a 12 GB card works with the compressed cache setting), plus PrismML's own llama.cpp build: as of late September, stock llama.cpp, Ollama and LM Studio can't load its main files. In practice, one developer's own test matrix passed 31 of 60 agentic coding tasks, and PrismML's own demo notes that only about half of its first attempts produced a game; the model thinks long by default and sometimes loops on tool calls. The free, private, one-prompt helper is here. The unsupervised agent replacement for your $10, 20 subscription is not. For builders weighing local compression against paid assistants, and anyone running sensitive code offline. Chapters: 0:00 Intro 1:20 Quick help 3:28 Agent work 6:15 In practice 8:52 Setup 11:18 Worth it? Tools & resources mentioned: - Ternary Bonsai 2 27B (Hugging Face): https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf - PrismML llama.cpp fork: https://github.com/PrismML-Eng/llama.cpp - PrismML Bonsai-demo repository: https://github.com/PrismML-Eng/Bonsai-demo - Bonsai 2 whitepaper: https://github.com/PrismML-Eng/Bonsai-demo/blob/main/bonsai-2-27b-whitepaper.pdf - PrismML launch post: https://prismml.com/news/bonsai-2-27b - SWE-bench Verified: https://www.swebench.com/verified.html - Terminal-Bench: https://www.tbench.ai/news/announcement - Hermes agent with Bonsai 2 (PrismML docs): https://docs.prismml.com/integrations/hermes - Cline coding tool - LiveCodeBench v6 About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #bonsai 2 #free ai coding #local models