How Fast Can Local AI Get On A 128GB Mini PC?

From the creator

A 128GB Ryzen AI Max+ 395 mini PC can write text fast, but the wait before the first word is a separate number. Here is what the published tests actually measured. On a pre-production Framework Desktop (Ryzen AI Max+ 395, 128GB), developer lhl's published llama.cpp sweep measured Qwen3 30B-A3B, a mixture-of-experts model, writing at 78.0 tokens per second, against 51.2 for the much smaller dense Llama 2 7B. In the same table Llama 2 7B read a short prompt faster: 924.9 tokens per second against 645.5. Reading the prompt and writing the answer are two different speeds, and llama-bench measures them as separate tests. Model size: Qwen3 30B-A3B holds 30.5 billion parameters but activates about 3.3 billion per token (8 of its 128 experts), which is why it writes quickly. Dense models use every parameter for every token: in the same sweep Mistral Small 3.1 (24B) wrote at 14.4 tokens per second and Shisa V2 70B at 5.0. All of them fit in 128GB, but fitting is not the same as being fast, and that memory is shared with the operating system. Long prompts: in a separate report on a retail Bosgame M5, with a different model (Qwen3.8 Flash Next) and a custom tuned llama.cpp build, developer drluoto measured 44 seconds to the first token on a fresh 18,000-token prompt, down from 67 seconds before tuning; the median writing speed across ten replayed conversations was about 33 tokens per second. Reading slowed as prompts grew, from 517 tokens per second at 8K to 391 at 32K, so a 512-token test cannot be stretched to a long document. Caching skips re-reading earlier context only while that cached state is still available. Tuning: the same report's speculative decoding branch, where a small draft head proposes tokens and the main model checks them, lifted a file rewrite job from 29.0 to 55.4 tokens per second (63.2 with six draft tokens) while prose barely moved (29.3 to 30.1), and the six-token setting halves prose speed. It needs the author's branch: stock llama.cpp will not load the draft file. Verdict: start with an efficient model that clears your quality bar with memory headroom to spare, with small-active mixture-of-experts models as strong candidates when you need fast, long answers. A heavier model can be worth the wait for complex reasoning, and heavy fresh documents call for measuring the first wait. Before declaring a winner, time your own task from prompt to usable answer: the first token and the writing rate. The Stack did not run these tests; the figures belong to lhl and drluoto. Sources and media credits: Benchmark results: lhl, Strix Halo LLM Benchmark Results (GitHub); drluoto, llama.cpp discussion #28512 and branch README (GitHub) Product footage and images: Framework, BOSGAME, AMD Stock footage (Pexels): Yan Krukau Tux by Larry Ewing, created with The GIMP Chapters: 0:00 Intro 1:00 Two speeds 2:59 Model size 5:04 Long prompts 7:43 Tuning 9:47 Your box 12:02 Conclusion Tools & resources mentioned: - lhl: Strix Halo LLM Benchmark Results: https://github.com/lhl/strix-halo-testing/tree/main/llm-bench - llama.cpp llama-bench README: https://github.com/ggml-org/llama.cpp/blob/master/tools/llama-bench/README.md - Qwen3-30B-A3B model card: https://huggingface.co/Qwen/Qwen3-30B-A3B - drluoto: Qwen3.8-Flash-Next on Strix Halo (llama.cpp discussion #28512): https://github.com/ggml-org/llama.cpp/discussions/28512 - drluoto: strix-halo-vulkan branch README: https://github.com/drluoto/llama.cpp/blob/strix-halo-vulkan/docs/strix-halo-flash-next-vulkan.md - llama.cpp: https://github.com/ggml-org/llama.cpp - Framework Desktop: https://frame.work - Bosgame M5 mini PC About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #localai #llm #mixtureofexperts

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.