Nemotron 3 Ultra: Is NVIDIA a Model Company Now?
NVIDIA just released Nemotron 3 Ultra, a 550B mixture-of-experts model built on a hybrid Transformer-Mamba architecture, and it's a clear sign of how far NVIDIA has moved beyond being just a hardware company. In this video I break down what the model is actually good at and where it still lags, the other open-weight models NVIDIA is shipping across speech, retrieval, robotics, and world models, and the business logic behind giving it all away for free. I'll also walk through how to access Nemotron 3 Ultra via NVIDIA's API, including thinking, reasoning budgets, and tool calling. Thanks to @NVIDIADeveloper for early access. Blog: https://nvda.ws/3PTkjlQ Hugging Face: https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 Tech Report: https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Ultra-Technical-Report.pdf Cookbook: https://github.com/NVIDIA-NeMo/Nemotron/tree/main/usage-cookbook/Nemotron-3-Ultra/ My voice to text App: whryte.com Website: https://engineerprompt.ai/ RAG Beyond Basics Course: https://prompt-s-site.thinkific.com/courses/rag Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 Let's Connect: 🦾 Discord: https://discord.com/invite/t4eYQRUcXB ☕ Buy me a Coffee: https://ko-fi.com/promptengineering |🔴 Patreon: https://www.patreon.com/PromptEngineering 💼Consulting: https://calendly.com/engineerprompt/consulting-call 📧 Business Contact: engineerprompt@gmail.com Become Member: http://tinyurl.com/y5h28s6h 💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off). Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0