What A 128GB Mini PC Can Actually Run

From the creator

Bosgame M5 mini PC with 128GB unified memory: can a $2,999 local AI desktop actually run 120B models faster than an RTX 4090? The Bosgame M5 packs AMD's Ryzen AI Max+ 395 with 128GB of unified LPDDR5X memory shared between CPU and GPU, enough space for massive open-weight models that won't fit on any 24GB consumer graphics card. But fitting a model and running it well are completely different problems. This video walks through which models actually deliver conversational speed on the M5 (mixture-of-experts architectures do; dense 70B models crawl at 4, 5 tok/s), why AMD's 2.2x faster than RTX 4090 claim is only true when the 4090 must spill into system RAM, and what real-world performance looks like across llama.cpp, ModelFit estimates, and MindStudio benchmarks. You'll learn why prompt processing becomes a brutal compute bottleneck for long-context coding tasks, how much driver surgery and compiler flag tuning the setup actually requires, why the $1,699 launch price jumped to $2,999 today, and exactly which buyer this machine is built for. Built for builders evaluating local AI hardware seriously, not hype seekers or those expecting plug-and-play convenience. Chapters: 0:00 Intro 1:07 Basics 2:48 Models 5:26 Speed 7:20 Prompts 8:55 Setup 10:57 Price 12:43 Worth it? 15:01 Conclusion Tools & resources mentioned: - Bosgame M5: https://www.bosgame.com/products/bosgame-m5-ai-mini-desktop-ryzen-ai-max-395-96gb-128gb-2tb - llama.cpp - ModelFit.io: https://modelfit.io/gpu/ryzen-ai-max-395 - MindStudio: https://www.mindstudio.ai/blog/amd-ryzen-ai-max-plus-395-local-ai-workstation - AMD ROCm - Qwen 3.8 Flash Next - GPT-OSS-120B - Open WebUI About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #local ai #open webui #llm

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.