NVIDIA's 84GB Card Ends The VRAM Ceiling

From the creator

NVIDIA's RTX PRO 5500 Blackwell packs 84GB VRAM to run 100B-param AI locally on one card, here's what actually fits, how fast, and who should buy. NVIDIA has quietly listed the RTX PRO 5500 Blackwell Workstation Edition with 84GB of GDDR7 memory, the same 21,760 CUDA cores as the GeForce RTX 5090, but positioned between the RTX PRO 5000 and RTX PRO 6000 in the professional stack. This extra memory lets you run large open models like OpenAI's GPT-OSS-120B whole on one card instead of splitting across several, a critical difference: in llama.cpp's gpt-oss guide, a 32GB RTX 5090 that has to keep part of the model in system RAM writes about 30 tokens/second, while an RTX PRO 6000 that holds all of it writes about 196. The spec sheet shows the 5500 should land between its two tested siblings in speed on models that fit entirely in graphics memory, driven by its ~1,398 GB/s memory bandwidth rather than core count. However, nobody has published real benchmarks yet; NVIDIA's page shows 'Coming Soon' with no price or ship date announced. The video breaks down which models actually fit (Llama 3.3 70B at four bits: yes; DeepSeek V4 Flash: barely; GLM-5.3: no), why memory bandwidth matters more than CUDA cores for LLM inference, and why splitting oversized models across multiple 5090s mostly buys room rather than speed (the link between the cards can become the bottleneck), and why three 5090s (needed to match 84GB) pull nearly three times the power of one 5500. NVIDIA pitched this card for IT departments to rack-mount and share across teams via Multi-Instance GPU splitting into two 42GB halves, not for individual desktop upgrades. This breakdown is for builders considering single-card local LLM serving, those evaluating multi-GPU workstations, and anyone curious how NVIDIA's professional AI stack actually maps to real model inference. Chapters: 0:00 Intro 1:28 What fits 4:02 Speed 6:41 Multi-GPU 9:02 Worth it? 10:49 Conclusion Tools & resources mentioned: - llama.cpp: https://github.com/ggerganov/llama.cpp - NVIDIA RTX PRO 5500 Blackwell: https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-5500 About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #nvidia #local llm #ai #gpu #rtx

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.