120B AI Models Beat 70B Ones On AMD's Mini PC

From the creator

Qwen 3.8 MoE blows past dense models on AMD Strix Halo, here's why bigger LLMs actually generate faster, and which token speeds to expect in llama.cpp. On AMD's Ryzen AI Max+ 395 with 128GB unified memory, a 109-billion-parameter mixture-of-experts model generates tokens faster than a 24-billion dense model, counterintuitively proving that on bandwidth-limited hardware, active parameters matter far more than total model size. Qwen 3.8 Flash Next (125B total, 6B active) and GPT-OSS-120B (117B total, 5.1B active) hit 15, 56 tokens/sec in llama.cpp depending on backend (Vulkan vs ROCm) and layer offloading, while dense seventy-billion models crawl at 4.5, 5 tokens/sec on the same silicon. The benchmark sweep from Level1Techs, Soot/Silicon tuning runs, and Framework community results show that Qwen3-30B-A3B MoE reaches 72 tok/s and Qwen3.6-35B-A3B achieves 42, 51 tok/s (Vulkan), demolishing the dense ladder. Backend choice and compilation flags (ROCBLAS_USE_HIPBLASLT, GGML_HIP_ROCWMMA_FATTN) swing performance by 6.4x on identical hardware; stock ROCm versus tuned Vulkan can mean 1.79 versus 15.56 tokens/sec on Qwen 3.8. The practical rule: ignore headline parameter count and chase the active parameters column instead, because every word requires the chip to read only the active slice, not the full model. For builders and AI-curious engineers evaluating Strix Halo mini PCs from Framework, GMKtec, or Beelink, and deciding which open-weights models to actually run locally. Chapters: 0:00 A 109B model outruns a 24B, here's why 0:27 The 128GB promise that sells Strix Halo 1:46 Where capacity meets the memory wall 3:07 Why every word reads the whole model 4:52 The dense speedrun ends at 24B 6:16 Giant models, tiny active slice 7:45 What actually matters: active parameters 9:55 Software tuning rewrites the hardware story 12:03 The real rule for your shopping cart Tools & resources mentioned: - llama.cpp: https://github.com/ggerganov/llama.cpp - LM Studio: https://lmstudio.ai - Open WebUI: https://openwebui.com - Qwen 3.8 Flash Next: https://huggingface.co/Qwen/Qwen3.8-Flash-Next - GPT-OSS-120B: https://huggingface.co/openai/gpt-oss-120b About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #Qwen3.8 #LocalLLM #StrixHalo

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.