Your Mac Might Be The Better Local AI Box
Ternary Bonsai 2 27B compresses Qwen 3.8 to 5.95 GB, runs on 8GB cards at 11, 40 tokens/sec, but barely on 6GB. Here's the real memory math. Ternary Bonsai 2 27B is PrismML's ternary-quantized compression of Alibaba's Qwen 3.8-27B: every weight is, 1, 0, or +1 with a shared scale per 128 weights, shrinking the model from 54 GB to 5.95 GB (PTQ1_0 file) while retaining 98.2% of the full model's benchmark average. On a 6 GB card, the file alone leaves ~0.5 GB free, but running it also demands a KV cache (64 KiB per token at default precision, 18 KiB with 4-bit cache), a fixed state for 48 linear-attention layers (~150 MiB), working buffers, and driver overhead, so only 44 of 65 layers fit on-card, producing 0.67 tokens/sec with CPU offload. At 8 GB, with tighter settings (smaller batch, compressed cache), the entire model runs on-card: RTX 3070 at 40.6 tokens/sec (32K context, 8-bit cache), RTX 2060 SUPER at 11 tokens/sec long-form (32K context). At 12 GB and above, the full 262K context fits comfortably at 26, 40 tokens/sec; Macs run it at 7, 8 tokens/sec (M2) to 47 tokens/sec (M5 Max) from shared memory. The model requires PrismML's own llama.cpp build, stock llama.cpp rejects the files or produces garbage due to Hadamard-rotated weights. LM Studio and Ollama cannot load it. On coding: short tests match the full model, but agentic tasks (SWE-bench Verified 60.8 vs. 80.6) retain ~75% of the original. Default xhigh reasoning loops excessively; medium avoids it. For 8 GB card owners running small coding tasks, chat, and document Q&A, this is the only way to run 27B entirely on-card. Chapters: 0:00 Intro 1:11 Memory 4:14 8GB Cards 7:34 Bigger Machines 10:10 Setup 11:39 Coding 15:27 Worth it? Tools & resources mentioned: - Ternary Bonsai 2 27B (PrismML): https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf - PrismML llama.cpp fork: https://github.com/PrismML-Eng/llama.cpp - Bonsai-demo (PrismML): https://github.com/PrismML-Eng/Bonsai-demo - Qwen 3.8-27B: https://huggingface.co/Qwen/Qwen3.8-27B About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #qwen38 #edgeai #gguf