Your Mac Might Be The Better Local AI Box

From the creator

Ternary Bonsai 2 27B compresses Qwen 3.8 to 5.95 GB, runs on 8GB cards at 11, 40 tokens/sec, but barely on 6GB. Here's the real memory math. Ternary Bonsai 2 27B is PrismML's ternary-quantized compression of Alibaba's Qwen 3.8-27B: every weight is, 1, 0, or +1 with a shared scale per 128 weights, shrinking the model from 54 GB to 5.95 GB (PTQ1_0 file) while retaining 98.2% of the full model's benchmark average. On a 6 GB card, the file alone leaves ~0.5 GB free, but running it also demands a KV cache (64 KiB per token at default precision, 18 KiB with 4-bit cache), a fixed state for 48 linear-attention layers (~150 MiB), working buffers, and driver overhead, so only 44 of 65 layers fit on-card, producing 0.67 tokens/sec with CPU offload. At 8 GB, with tighter settings (smaller batch, compressed cache), the entire model runs on-card: RTX 3070 at 40.6 tokens/sec (32K context, 8-bit cache), RTX 2060 SUPER at 11 tokens/sec long-form (32K context). At 12 GB and above, the full 262K context fits comfortably at 26, 40 tokens/sec; Macs run it at 7, 8 tokens/sec (M2) to 47 tokens/sec (M5 Max) from shared memory. The model requires PrismML's own llama.cpp build, stock llama.cpp rejects the files or produces garbage due to Hadamard-rotated weights. LM Studio and Ollama cannot load it. On coding: short tests match the full model, but agentic tasks (SWE-bench Verified 60.8 vs. 80.6) retain ~75% of the original. Default xhigh reasoning loops excessively; medium avoids it. For 8 GB card owners running small coding tasks, chat, and document Q&A, this is the only way to run 27B entirely on-card. Chapters: 0:00 Intro 1:11 Memory 4:14 8GB Cards 7:34 Bigger Machines 10:10 Setup 11:39 Coding 15:27 Worth it? Tools & resources mentioned: - Ternary Bonsai 2 27B (PrismML): https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf - PrismML llama.cpp fork: https://github.com/PrismML-Eng/llama.cpp - Bonsai-demo (PrismML): https://github.com/PrismML-Eng/Bonsai-demo - Qwen 3.8-27B: https://huggingface.co/Qwen/Qwen3.8-27B About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #qwen38 #edgeai #gguf

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.