Your 128GB Of RAM Won't Save A 24GB Card For Local AI
Local AI on 24GB GPU vs 64GB Mac? Why prompt reading speed decides this build, measured against real benchmarks. Running Qwen's 35B MoE model as a long-context AI agent means spending most of your time reading prompts, not writing answers, and that's a compute problem, not a memory one. A Reddit buyer on a 32GB M2 Max waited 20 minutes to process a 40,000-token prompt; the question became whether a used RTX 3090 with 128GB RAM or a 64GB Mac would run it faster. In llama.cpp, both machines read short prompts at ~3,100 tokens per second, but at 8,192 tokens the RTX 3090 held 2,965 tok/s while the M5 Max dropped to 2,036. The 3090 generates ~143 tok/s; the M5 Max ~85. Neither machine needs that 128GB of system memory for this model: the full 20.6GB file fits on the 24GB card with room for ~90,000 tokens of context before speed degrades, while a 64GB Mac holds the entire 262,144-token context natively. The real cost reveal: 128GB of DDR5 RAM now costs $1,950, more than a used 3090 (~$1,400, 1,500). Apple's 64GB Mac Studio with M5 Max is $3,799; the M5 Pro Mac mini 64GB is $3,199 but reads prompts at half the M5 Max speed. For agentic workflows where prompt latency kills productivity, this video breaks down which hardware actually delivers. Chapters: 0:00 Intro 2:11 Reading 3:52 Speed 6:24 Fit 9:22 Memory 11:39 Verdict Tools & resources mentioned: - llama.cpp: https://github.com/ggml-org/llama.cpp - Qwen3.6-35B-A3B model card: https://huggingface.co/Qwen/Qwen3.6-35B-A3B - Apple MLX research: https://machinelearning.apple.com/research/exploring-llms-mlx-m5 - Mac Studio 64GB M5 Max: https://www.apple.com/shop/buy-mac/mac-studio/m5-max-chip-18-core-cpu-40-core-gpu-64gb-memory-1tb-storage - Mac mini 64GB M5 Pro: https://www.apple.com/shop/buy-mac/mac-mini/m5-pro-chip-18-core-cpu-20-core-gpu-64gb-memory-1tb-storage About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #local-ai #llama-cpp #qwen-moe #ai-agents #apple-silicon