What A 128GB Mini PC Can Actually Run
Bosgame M5 mini PC with 128GB unified memory: can a $2,999 local AI desktop actually run 120B models faster than an RTX 4090? The Bosgame M5 packs AMD's Ryzen AI Max+ 395 with 128GB of unified LPDDR5X memory shared between CPU and GPU, enough space for massive open-weight models that won't fit on any 24GB consumer graphics card. But fitting a model and running it well are completely different problems. This video walks through which models actually deliver conversational speed on the M5 (mixture-of-experts architectures do; dense 70B models crawl at 4, 5 tok/s), why AMD's 2.2x faster than RTX 4090 claim is only true when the 4090 must spill into system RAM, and what real-world performance looks like across llama.cpp, ModelFit estimates, and MindStudio benchmarks. You'll learn why prompt processing becomes a brutal compute bottleneck for long-context coding tasks, how much driver surgery and compiler flag tuning the setup actually requires, why the $1,699 launch price jumped to $2,999 today, and exactly which buyer this machine is built for. Built for builders evaluating local AI hardware seriously, not hype seekers or those expecting plug-and-play convenience. Chapters: 0:00 Intro 1:07 Basics 2:48 Models 5:26 Speed 7:20 Prompts 8:55 Setup 10:57 Price 12:43 Worth it? 15:01 Conclusion Tools & resources mentioned: - Bosgame M5: https://www.bosgame.com/products/bosgame-m5-ai-mini-desktop-ryzen-ai-max-395-96gb-128gb-2tb - llama.cpp - ModelFit.io: https://modelfit.io/gpu/ryzen-ai-max-395 - MindStudio: https://www.mindstudio.ai/blog/amd-ryzen-ai-max-plus-395-local-ai-workstation - AMD ROCm - Qwen 3.8 Flash Next - GPT-OSS-120B - Open WebUI About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #local ai #open webui #llm