NVIDIA’s 4-Bit Format Makes Local AI 2X Faster On Your GPU

From the creator

NVFP4 quantization on Blackwell GPUs speeds local AI inference, but not 2X. See real benchmarks, memory trade-offs, and whether it's worth converting your models. NVIDIA sold FP4 on its RTX 50 cards as "2X", but that figure came from image generation: FLUX running in FP4 on an RTX 5090 against FP16 on an RTX 4090, in NVIDIA's own comparison. NVFP4 is NVIDIA's 4-bit floating-point format. Values are stored as 4-bit numbers in blocks of 16 that share an FP8 scale, which reduces rounding error although some remains, and Blackwell's Tensor Cores can calculate with them directly when the runtime routes the work there (llama.cpp added native Blackwell NVFP4 kernels in April 2026). In a developer's before-and-after test of the same NVFP4 file of Qwen3.5-27B, native kernels made prompt processing about 46% faster but reply generation only about 1% faster, on unspecified hardware. LibertAI's own test of Qwen3.6-27B on an RTX 5090, one request at a time with 512 tokens in and 128 out, reported 12% faster generation and 9% more total throughput than standard Q4_K_M. Memory: their hybrid conversion, with NVFP4 only in the large feed-forward layers, used 14.7 GiB against 16.3 GiB, but a roughly 15 GB file does not prove a whole chat fits, because the runtime also needs workspace and the cache grows with longer context. Quality: a preliminary developer check on the small Qwen3.5-4B found stock Q4_K closer to the reference than NVFP4 on statistical metrics, and the 27B speed test included no matched quality evaluation. On a Blackwell card, compare conversions of the same model in one runtime on your real work, check the runtime's documentation and startup logs (a log showing every layer on the GPU does not prove native NVFP4 math), match your settings, and time the whole task while checking the answers against your source. On an older card, these results alone are no reason to buy a new GPU. Media credits: Footage: NVIDIA, CES 2025 keynote https://www.youtube.com/watch?v=k82RwXqZHY8 Footage: NVIDIA GeForce, GeForce RTX 5090 Founders Edition Overview https://www.youtube.com/watch?v=EZ5UBhEDm-I Product photos and the RTX 5090 performance chart: NVIDIA GeForce News, 6 Jan 2025 Stock footage (Pexels): SHVETS production, Tima Miroshnichenko llama.cpp logo: github.com/ggml-org/llama.brand Chapters: 0:00 Intro 1:00 The trick 2:50 Speed 5:16 Memory 6:54 Quality 8:26 Your GPU 10:42 Conclusion Tools & resources mentioned: - NVIDIA GeForce News: RTX 50 Series and the 2X chart: https://www.nvidia.com/en-us/geforce/news/rtx-50-series-graphics-cards-gpu-laptop-announcements/ - NVIDIA Technical Blog: Introducing NVFP4: https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/ - LibertAI Qwen3.6-27B NVFP4 GGUF and their benchmark: https://huggingface.co/LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF - llama.cpp PR #21896: native NVFP4 before and after: https://github.com/ggml-org/llama.cpp/pull/21896 - llama.cpp PR #22196: Blackwell NVFP4 kernels (merged): https://github.com/ggml-org/llama.cpp/pull/22196 - llama.cpp discussion #22498: NVFP4 vs Q4_K check: https://github.com/ggml-org/llama.cpp/discussions/22498 - llama.cpp: https://github.com/ggml-org/llama.cpp - llama-bench: https://github.com/ggml-org/llama.cpp/blob/master/tools/llama-bench/README.md - NVIDIA: LLM inference benchmarking concepts: https://developer.nvidia.com/blog/llm-benchmarking-fundamental-concepts/ About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #NVFP4 #LocalLLM #LLMQuantization

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.