NVIDIA's 84GB Card Ends The VRAM Ceiling
NVIDIA's RTX PRO 5500 Blackwell packs 84GB VRAM to run 100B-param AI locally on one card, here's what actually fits, how fast, and who should buy. NVIDIA has quietly listed the RTX PRO 5500 Blackwell Workstation Edition with 84GB of GDDR7 memory, the same 21,760 CUDA cores as the GeForce RTX 5090, but positioned between the RTX PRO 5000 and RTX PRO 6000 in the professional stack. This extra memory lets you run large open models like OpenAI's GPT-OSS-120B whole on one card instead of splitting across several, a critical difference: in llama.cpp's gpt-oss guide, a 32GB RTX 5090 that has to keep part of the model in system RAM writes about 30 tokens/second, while an RTX PRO 6000 that holds all of it writes about 196. The spec sheet shows the 5500 should land between its two tested siblings in speed on models that fit entirely in graphics memory, driven by its ~1,398 GB/s memory bandwidth rather than core count. However, nobody has published real benchmarks yet; NVIDIA's page shows 'Coming Soon' with no price or ship date announced. The video breaks down which models actually fit (Llama 3.3 70B at four bits: yes; DeepSeek V4 Flash: barely; GLM-5.3: no), why memory bandwidth matters more than CUDA cores for LLM inference, and why splitting oversized models across multiple 5090s mostly buys room rather than speed (the link between the cards can become the bottleneck), and why three 5090s (needed to match 84GB) pull nearly three times the power of one 5500. NVIDIA pitched this card for IT departments to rack-mount and share across teams via Multi-Instance GPU splitting into two 42GB halves, not for individual desktop upgrades. This breakdown is for builders considering single-card local LLM serving, those evaluating multi-GPU workstations, and anyone curious how NVIDIA's professional AI stack actually maps to real model inference. Chapters: 0:00 Intro 1:28 What fits 4:02 Speed 6:41 Multi-GPU 9:02 Worth it? 10:49 Conclusion Tools & resources mentioned: - llama.cpp: https://github.com/ggerganov/llama.cpp - NVIDIA RTX PRO 5500 Blackwell: https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-5500 About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #nvidia #local llm #ai #gpu #rtx