Your 128GB Of RAM Won't Save A 24GB Card For Local AI

From the creator

Local AI on 24GB GPU vs 64GB Mac? Why prompt reading speed decides this build, measured against real benchmarks. Running Qwen's 35B MoE model as a long-context AI agent means spending most of your time reading prompts, not writing answers, and that's a compute problem, not a memory one. A Reddit buyer on a 32GB M2 Max waited 20 minutes to process a 40,000-token prompt; the question became whether a used RTX 3090 with 128GB RAM or a 64GB Mac would run it faster. In llama.cpp, both machines read short prompts at ~3,100 tokens per second, but at 8,192 tokens the RTX 3090 held 2,965 tok/s while the M5 Max dropped to 2,036. The 3090 generates ~143 tok/s; the M5 Max ~85. Neither machine needs that 128GB of system memory for this model: the full 20.6GB file fits on the 24GB card with room for ~90,000 tokens of context before speed degrades, while a 64GB Mac holds the entire 262,144-token context natively. The real cost reveal: 128GB of DDR5 RAM now costs $1,950, more than a used 3090 (~$1,400, 1,500). Apple's 64GB Mac Studio with M5 Max is $3,799; the M5 Pro Mac mini 64GB is $3,199 but reads prompts at half the M5 Max speed. For agentic workflows where prompt latency kills productivity, this video breaks down which hardware actually delivers. Chapters: 0:00 Intro 2:11 Reading 3:52 Speed 6:24 Fit 9:22 Memory 11:39 Verdict Tools & resources mentioned: - llama.cpp: https://github.com/ggml-org/llama.cpp - Qwen3.6-35B-A3B model card: https://huggingface.co/Qwen/Qwen3.6-35B-A3B - Apple MLX research: https://machinelearning.apple.com/research/exploring-llms-mlx-m5 - Mac Studio 64GB M5 Max: https://www.apple.com/shop/buy-mac/mac-studio/m5-max-chip-18-core-cpu-40-core-gpu-64gb-memory-1tb-storage - Mac mini 64GB M5 Pro: https://www.apple.com/shop/buy-mac/mac-mini/m5-pro-chip-18-core-cpu-20-core-gpu-64gb-memory-1tb-storage About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #local-ai #llama-cpp #qwen-moe #ai-agents #apple-silicon

Choose to Build with AI
Matched to AI Agents

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.