Build a Local Router with Nemotron Lightning
Thanks to @NVIDIADeveloper for DGX Spark. Check it out here: https://nvda.ws/3XIkwsh In this video I explain why a routing layer is becoming essential for agentic systems and show how to run your own router locally so you control cost, speed, specialization, and privacy. I compare proprietary and open-source options (OpenRouter, Devin Fusion, RouteLLM) and then focus on NVIDIA Switchyard (built on RouteLLM) and how it supports multiple routing strategies: random, LLM classifier, stage routing, and escalation—plus their cost tradeoffs. I also cover NVIDIA Nemotron 3.5 Lightning (30B latent MoE) and why its high throughput matters, including a speed test versus Kimi K3. Finally, I demo local-first alert analysis workflows where Nemotron generates private digests and Switchyard escalates to Kimi K3 only when needed, and explain why routing should be done per task/session rather than per turn. LINKS: Blog: https://nvda.ws/3RAV34N Nemotron 3.5 Lightning: https://nvda.ws/3RPClGE NeMo Switchyard: https://nvda.ws/4wO4WLE NVFP4 DFlash: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash NVFP4 DSpark: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark NVFP4: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 My voice to text App: whryte.com Website: https://engineerprompt.ai/ RAG Beyond Basics Course: https://prompt-s-site.thinkific.com/courses/rag Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 00:00 Why Routers Matter 01:29 Cost Speed Privacy 03:09 Routing Options Today 04:01 Nemotron Lightning Intro 05:19 Speed Test Results 06:34 Switchyard Setup Basics 07:27 Four Routing Strategies 08:50 Signals And Cost Tradeoffs 09:28 Cutting Costs With Hybrid 10:33 Network Alerts Demo 12:22 Per Alert Escalation 12:56 Best Practices And Wrap