Run a 405B Model With OPNsense on a DGX Spark

From the creator

DGX Spark clustering: two NVIDIA GB10 boxes, one cable, 256GB unified memory to run a 405B local LLM at home Running a 405B-parameter open model on a desk sounds impossible until you understand unified memory. This breaks down how two NVIDIA DGX Spark units, each built on the GB10 Grace Blackwell superchip with 128GB unified memory, link over a single QSFP56 cable and 200GbE ConnectX-7 port to pool 256GB and roughly 2 PFLOP FP4 compute, enough headroom to load a 405B-class model instead of the ~70B ceiling a single unit hits. It also covers why the same GB10 chip ships in rival boxes like the ASUS Ascent GX10, how Ollama and Tailscale make the software side a solved problem for local ai and local chatgpt style workflows, and why NVIDIA quietly raised the Spark's price from $3,999 to $4,699 in early 2026. The second half is about what guards that hardware once it's on your desk. OPNsense shows up as the open-source router/firewall of choice for home lab ai builders who ditched their ISP gateway, and the video walks through real 2026 CVEs, a 9.1-severity root RCE (CVE-2026-44194), a DHCP-triggered RCE (CVE-2026-45158), a CSRF flaw (CVE-2026-30868), and an auth-lockout bypass (CVE-2026-44195), to show what it actually costs to be your own network admin. Along the way it touches on gpu vram upgrade logic, rtx 4090 and rtx 4090 48gb builds, m3 ai comparisons, and where open webui fits into a self-hosted stack. For anyone weighing an ai server build against another month of cloud API bills, this lays out the real trade-off: capacity versus speed, and control versus who patches your front door. Chapters: 0:00 The Model Nobody Can Actually Hold 1:40 What A Desktop Supercomputer Costs You 3:02 The Secret Chip Inside Every Rival Box 4:28 The One Cable That Doubles Your Memory 5:47 Is This Machine Actually Worth It 6:58 The Risk You Bring Home With It 8:10 Building The Front Door You Own 9:38 The Trade Nobody Warns You About 10:45 The Week The Firewall Broke Open 12:25 The Hack That Needs No Password 13:53 Just How Bad Can It Get 15:32 So Can A Desk Really Hold It Tools & resources mentioned: - NVIDIA DGX Spark: https://www.nvidia.com/en-us/products/workstations/dgx-spark/ - OPNsense: https://opnsense.org/ - Ollama: https://ollama.com/ - Tailscale: https://tailscale.com/ - ASUS Ascent GX10: https://www.asus.com/ - OpenWrt: https://openwrt.org/ About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #dgxspark #localllm #homelabai #opnsense #selfhosted

Choose to Build with AI
Matched to Open WebUI

AI Maker Residence 3

The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence 3
Live event
AI Maker Residence 3
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.