Can Unsloth Turn An AMD GPU Into A Local AI Powerhouse?
Unsloth Desktop now runs open models and agents on AMD Radeon cards, here's what actually works, what crashes, and whether to use it today. Unsloth, long known for fine-tuning with lower memory overhead, now ships as a free desktop app that runs open models locally and connects them to coding agents like Claude Code. AMD GPU support officially launched in July 2026, and the platform claims to work across Radeon RX 6000/7000/9000 series cards, Ryzen AI Max chips, and Windows/WSL/Linux. This video tests the real state of that support using Qwen 3.8 27B (a ~17.5GB four-bit file) on hardware like the RX 7900 XTX, walking through hardware compatibility, how the OpenAI-compatible local API wires agents to your card, realistic speed measurements from LocalScore benchmarks (an RX 7900 XTX runs Qwen 14B at ~28 tokens/sec generation, ~500 tokens/sec for prompt reading), QLoRA training memory footprints, and a critical GitHub report of driver crashes during AMD training as of late September. You'll learn which older Radeon cards only support inference, why ROCm setup is automated but still beta-stage, how to avoid the 90% slowdown from Claude Code's per-request prompt resets, and whether small training experiments are safe. For builders deciding whether to run a local coding agent on AMD hardware today, this covers the working paths, the real gotchas, and the current stability picture. Chapters: 0:00 Intro 1:30 Hardware 3:05 Coding agents 4:50 Speed 6:50 Training 8:54 Verdict Tools & resources mentioned: - Unsloth Desktop: https://unsloth.ai - Qwen (Hugging Face): https://huggingface.co/Qwen - llama.cpp - Claude Code - Codex - LocalScore benchmarks: https://localscore.ai - ROCm (AMD): https://rocmdocs.amd.com - Hugging Face About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #unsloth #AMD GPU #local AI #coding agents #Claude Code