Apple Silicon Now DOMINATES Local AI Hardware (M5 Max vs RTX 5090)

From the creator

M5 Ultra Mac Studio vs RTX 5090: why unified memory beats VRAM for local LLM inference, not raw speed The M5 Ultra Mac Studio doesn't beat the RTX 5090 on speed, it beats it on what actually decides local AI: memory capacity. This breaks down why unified memory lets Apple Silicon run models NVIDIA's flagship physically cannot load, and where each machine actually wins. The core problem is that the RTX 5090 ships with a hard 32GB VRAM ceiling, which excludes any 70B-parameter model at Q4 quantization without CPU offload killing performance. Apple's unified memory architecture sidesteps that entirely, letting a Mac Studio load models like a 235B-parameter model that would otherwise require a multi-GPU NVIDIA cluster worth $30,000+. On raw throughput the 5090 still wins clean, three to four times faster where models overlap in size, thanks to nearly 1,800 GB/s of memory bandwidth versus the M4 Max's ~550 GB/s. But the video also covers why that speed gap stopped mattering as much: Ollama's switch to Apple's MLX framework nearly doubled decode speed on qualifying M4/M5 hardware, closing years of software disadvantage against CUDA. It covers total cost of ownership (a Mac Studio's power draw and resale value vs a 5090 rig), where MLX and Open WebUI or LM Studio fit into a local LLM stack, the ollama vs vllm question on Apple Silicon, and where NVIDIA and tools like DGX Spark still win outright, fine-tuning, image generation, and low-latency serving. For solo builders and small teams deciding between Apple Silicon and NVIDIA for local AI, machine learning experiments, or running an LLM at home without a five-figure GPU cluster, this lays out exactly which hardware tier fits which workload. Chapters: 0:00 Intro 0:14 The one spec that decides everything 2:01 Where NVIDIA still crushes it 3:01 The cost gap nobody expects 3:57 The software update that changed it all 5:32 NVIDIA's real wins, Apple's real hole 7:08 Which one should you actually buy Tools & resources mentioned: - MLX: https://ollama.com/blog/mlx - Ollama - Open WebUI - LM Studio - vLLM - RTX 5090 - DGX Spark - Mac Studio (M5 Ultra) - Mac Mini M4 Pro About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #localai #appleSilicon #macstudio #localllm #machinelearning

Choose to Build with AI
Matched to Open WebUI

AI Maker Residence 3

The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence 3
Live event
AI Maker Residence 3
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.