NVIDIA’s 4-Bit Format Makes Local AI 2X Faster On Your GPU
NVFP4 quantization on Blackwell GPUs speeds local AI inference, but not 2X. See real benchmarks, memory trade-offs, and whether it's worth converting your models. NVIDIA sold FP4 on its RTX 50 cards as "2X", but that figure came from image generation: FLUX running in FP4 on an RTX 5090 against FP16 on an RTX 4090, in NVIDIA's own comparison. NVFP4 is NVIDIA's 4-bit floating-point format. Values are stored as 4-bit numbers in blocks of 16 that share an FP8 scale, which reduces rounding error although some remains, and Blackwell's Tensor Cores can calculate with them directly when the runtime routes the work there (llama.cpp added native Blackwell NVFP4 kernels in April 2026). In a developer's before-and-after test of the same NVFP4 file of Qwen3.5-27B, native kernels made prompt processing about 46% faster but reply generation only about 1% faster, on unspecified hardware. LibertAI's own test of Qwen3.6-27B on an RTX 5090, one request at a time with 512 tokens in and 128 out, reported 12% faster generation and 9% more total throughput than standard Q4_K_M. Memory: their hybrid conversion, with NVFP4 only in the large feed-forward layers, used 14.7 GiB against 16.3 GiB, but a roughly 15 GB file does not prove a whole chat fits, because the runtime also needs workspace and the cache grows with longer context. Quality: a preliminary developer check on the small Qwen3.5-4B found stock Q4_K closer to the reference than NVFP4 on statistical metrics, and the 27B speed test included no matched quality evaluation. On a Blackwell card, compare conversions of the same model in one runtime on your real work, check the runtime's documentation and startup logs (a log showing every layer on the GPU does not prove native NVFP4 math), match your settings, and time the whole task while checking the answers against your source. On an older card, these results alone are no reason to buy a new GPU. Media credits: Footage: NVIDIA, CES 2025 keynote https://www.youtube.com/watch?v=k82RwXqZHY8 Footage: NVIDIA GeForce, GeForce RTX 5090 Founders Edition Overview https://www.youtube.com/watch?v=EZ5UBhEDm-I Product photos and the RTX 5090 performance chart: NVIDIA GeForce News, 6 Jan 2025 Stock footage (Pexels): SHVETS production, Tima Miroshnichenko llama.cpp logo: github.com/ggml-org/llama.brand Chapters: 0:00 Intro 1:00 The trick 2:50 Speed 5:16 Memory 6:54 Quality 8:26 Your GPU 10:42 Conclusion Tools & resources mentioned: - NVIDIA GeForce News: RTX 50 Series and the 2X chart: https://www.nvidia.com/en-us/geforce/news/rtx-50-series-graphics-cards-gpu-laptop-announcements/ - NVIDIA Technical Blog: Introducing NVFP4: https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/ - LibertAI Qwen3.6-27B NVFP4 GGUF and their benchmark: https://huggingface.co/LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF - llama.cpp PR #21896: native NVFP4 before and after: https://github.com/ggml-org/llama.cpp/pull/21896 - llama.cpp PR #22196: Blackwell NVFP4 kernels (merged): https://github.com/ggml-org/llama.cpp/pull/22196 - llama.cpp discussion #22498: NVFP4 vs Q4_K check: https://github.com/ggml-org/llama.cpp/discussions/22498 - llama.cpp: https://github.com/ggml-org/llama.cpp - llama-bench: https://github.com/ggml-org/llama.cpp/blob/master/tools/llama-bench/README.md - NVIDIA: LLM inference benchmarking concepts: https://developer.nvidia.com/blog/llm-benchmarking-fundamental-concepts/ About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #NVFP4 #LocalLLM #LLMQuantization