AMD's 128GB Ram Desktop Makes Local AI INSANELY Good
AMD's 128GB unified-memory desktop runs LM Studio and ComfyUI AMD workloads to find where Radeon 8060S actually holds up AMD's Ryzen AI Halo is a $3,999 desktop with no discrete GPU, running a Radeon 8060S iGPU and 128GB of unified LPDDR5x memory instead. This video runs it through four real local-AI workloads to see where that unified-memory design actually pays off and where it falls apart, testing LM Studio inference, fine-tuning, ComfyUI AMD image generation, and AI video generation on the same box. On chat inference, the machine holds a 120B parameter model that would OOM on any consumer GPU, hitting 45-53 tokens/sec in LM Studio and climbing to 66-97 t/s on smaller Qwen3 and Granite models, with independent llama.cpp benchmarks showing the ROCm HIP backend roughly doubling CPU-only speeds. Fine-tuning works too, but only through a community GitHub repo (amd-strix-halo-fine-tuning-toolboxes) since AMD never shipped official tooling. ComfyUI ROCm image generation runs but the iGPU only hits about 60% of theoretical peak, and Wan 2.2 video generation causes black screens and full reboots, taking roughly four times longer than an RTX 5090 when it does complete. A documented ROCm regression also broke stability for two months, requiring a specific Linux kernel and ROCm nightly pairing to fix. The video compares this against NVIDIA's DGX Spark, noting the AMD box's native Windows 11 dual-boot as a real advantage, and lays out a clear buying rule based on the results: unified memory wins when the problem is model size, not raw throughput. For anyone weighing a Radeon 8060S or AMD ROCm setup against a DGX Spark or discrete GPU for local AI, this breaks down exactly which workloads fit and which ones don't. Chapters: 0:00 The Desktop With No Graphics Card 1:42 The One Job It Wins Outright 3:09 Nobody Built The Software For This 4:47 Where The Cracks Start Showing 6:22 Two Months It Just Didn't Work 7:54 So Who Should Actually Buy This Tools & resources mentioned: - ComfyUI - LM Studio: https://lmstudio.ai - Open WebUI: https://openwebui.com - llama.cpp: https://github.com/ggerganov/llama.cpp - AMD ROCm: https://www.amd.com/en/products/software/rocm.html - Ollama: https://ollama.com - amd-strix-halo-fine-tuning-toolboxes: https://github.com/shantur/amd-strix-halo-fine-tuning-toolboxes - NVIDIA DGX Spark: https://www.nvidia.com/en-us/products/workstations/dgx-spark/ About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #comfyuiamd #localai #amdrocm #lmstudio #aiworkstation