A vendor's 4-bit throughput multiplier hides two decisions: which GPU generation you are committing to, and which capability your model loses. For most enterprise inference, run FP8 on Hopper or Blackwell: in an independent 16-arm benchmark it moved MMLU-Pro by 0.7 points or less on every model tested. When you need 4-bit to fit a bigger model or a longer context, use INT4-AWQ on mature Marlin kernels: on a Blackwell workstation card it decoded faster than NVFP4 on three of four models and produced the smaller file on both dense ones. NVFP4 loses on dense models today. It earns its place on mixture-of-experts models on Blackwell, ideally from a checkpoint NVIDIA recovered with quantization-aware distillation. GGUF belongs on laptops, CPUs and Macs, not in a GPU serving fleet.
| Format | Runs natively on | MMLU-Pro change vs BF16 (4 models) | Decode, Qwen3.6-27B, 1 stream | File size, Gemma-4-31B | Pick it when | Skip it when |
|---|---|---|---|---|---|---|
| FP8 (E4M3) | Ada, Hopper, Blackwell, AMD GPU | -0.1 to -0.7 | 48.8 tok/s | 33.3 GB | You have the memory. This is the default. | You are on Ampere or older |
| INT4-AWQ | Turing through Hopper in vLLM's table, plus Blackwell, Intel GPU, x86 CPU | -0.1 to -1.5 | 69.0 tok/s | 20.5 GB | You need 4-bit on a dense model, on any recent NVIDIA card | You serve at high batch on Blackwell and have measured NVFP4 winning there |
| NVFP4 | Blackwell only | +0.2 to -2.4 | 55.6 tok/s | 32.7 GB | MoE models on Blackwell, QAD checkpoints | Your fleet includes Hopper, or the model is small and dense |
| GGUF (llama.cpp) | CPU, Apple silicon, most GPUs | Not measured in this benchmark | Not measured | Not measured | Laptops, edge boxes, air-gapped desktops | You run a multi-user GPU service |
The workload behind the numbers: one RTX PRO 6000 Blackwell Max-Q (96 GB) running vLLM 0.22.1, single stream, 64k context, published June 8, 2026 by Stefano Schotten. Its code is MIT-licensed with both full runs committed under results/ and results-rerun/, and the checkpoints are pinned by SHA. That makes it the only cross-format comparison we found that you can rerun yourself.
What Each Format Is, and the GPU It Commits You To
Quantization is storing a model's weights, and sometimes its activations, in fewer bits than the 16 it was trained in. Each format is tied to specific silicon, so picking one is a hardware decision.
FP8 is an 8-bit floating-point format. NVIDIA added FP8 Tensor Cores with the H100, in two variants: E4M3 (4 exponent bits, 3 mantissa bits) for precision, and E5M2 for range. Inference uses E4M3. In vLLM's hardware table, full W8A8 FP8 is supported on Ada and Hopper among NVIDIA columns, plus AMD GPUs. Ampere only gets FP8 weights through the Marlin weight-only path.
NVFP4 is NVIDIA's 4-bit float for Blackwell. Each block of 16 values shares an FP8 E4M3 scale, with a second FP32 scale per tensor, and NVIDIA's fifth-generation Tensor Cores execute it in hardware. FP4 hardware support arrived with Blackwell, so an NVFP4 checkpoint commits you to Blackwell.
INT4-AWQ is 4-bit integer weights with 16-bit activations. AWQ finds the roughly 1% of weights that matter most from activation statistics and scales them up before rounding. GPTQ is the older alternative that compensates rounding error layer by layer. vLLM supports AWQ from Turing through Hopper and on Intel GPUs and x86 CPUs, which makes it the most portable 4-bit option. In a study of models from 1B to 405B parameters, AWQ tended to beat GPTQ on accuracy.
AutoRound is a recipe that emits these formats, and you should evaluate it as one. Intel's Apache-2.0 toolkit exports to AutoGPTQ, AutoAWQ, GGUF and llm-compressor, and supports NVFP4 and MXFP4 schemes.
GGUF is llama.cpp's file format, built for CPUs and consumer hardware. Its popular Q4_K_M preset mixes precisions and lands near 4.9 bits per weight on Llama-3.1-8B, above the "4-bit" label.
The hardware commitment has a price. On Lambda's on-demand list, checked October 9, 2026, an 8-GPU H100 SXM instance is $3.99 per GPU-hour and an 8-GPU B200 instance is $6.69 per GPU-hour, 68% more. A format that only runs natively on Blackwell has to pay for that gap in throughput you actually measure.
Where the Accuracy Loss Actually Lands
Four-bit costs little on average, and nearly all of that cost lands in knowledge recall. Averaged over MMLU-Pro, GSM8K, IFEval, HumanEval and MBPP, NVFP4 cost at most 0.6 points on the dense models in the independent benchmark. Math, code and instruction-following sat near their ceilings.
MMLU-Pro is where it shows. On Qwen3.6-27B, MMLU-Pro fell 0.4 points under FP8, 1.5 under INT4-AWQ and 2.4 under NVFP4. On Gemma-4-31B the drops were 0.7, 1.0 and 1.6. The two MoE models barely moved: Qwen3.6-35B-A3B lost 0.5 under NVFP4, and Gemma-4-26B-A4B gained 0.2. The author's reading is that NVFP4 and INT4-AWQ are a wash at equal bytes per parameter, and that the recipe decides the result more than the number format does.
NVIDIA's own numbers point the same way at larger scale. On DeepSeek-R1-0528, NVFP4 scored within one point of FP8 on MMLU-Pro (84% vs 85%), GPQA Diamond and LiveCodeBench, and two points higher on AIME 2024.
Two caveats change how you read this. The benchmark's retest found a mean drift of 0.35 points across 80 scores, with MMLU-Pro moving up to 1.7 points between identical runs, so sub-point gaps are inside the noise. And when the author checked NVIDIA's published deltas, they matched on Qwen3.6-35B-A3B (-0.57 measured vs -0.60 published) and diverged on Gemma-4-31B: -1.57 measured against NVIDIA's -0.31.
Why an NVFP4 File Can Be Bigger Than an AWQ File
NVFP4's scales cost real bits, and NVIDIA's checkpoints leave some layers at full precision, so "4-bit" files vary widely in size. NVIDIA puts NVFP4 at about 4.5 bits per value, roughly 1.8x smaller than FP8.
The files on disk tell a different story. In the benchmark's size table, Gemma-4-31B came to 20.5 GB as INT4-AWQ and 32.7 GB as NVIDIA's official NVFP4, within 0.6 GB of the 33.3 GB FP8 file. NVIDIA's recipe keeps most attention layers in BF16. Qwen3.6-27B was 21.9 GB as AWQ and 26.4 GB as the community NVFP4. On Gemma-4-26B-A4B, AWQ was again smaller (17.2 GB against 18.8 GB). NVFP4 produced the smallest file on one model of four, Qwen3.6-35B-A3B, at 23.5 GB against AWQ's 25.5 GB.
If you size GPU memory from the nominal bit width, you will be wrong by up to 12 GB on a 31B model. Size it from the checkpoint you will actually deploy.
Throughput Follows Kernel Maturity, Not the Bit Width
The format with the better on-paper spec lost most of the speed tests, because INT4-AWQ runs on the mature Marlin kernels. Single-stream decode on the Blackwell card, per the benchmark:
| Model | FP8 | INT4-AWQ | NVFP4 |
|---|---|---|---|
| Qwen3.6-27B (dense) | 48.8 | 69.0 | 55.6 |
| Gemma-4-31B-it (dense) | 42.0 | 63.9 | 40.8 |
| Qwen3.6-35B-A3B (MoE) | 221.4 | 194.2 | 226.3 |
| Gemma-4-26B-A4B (MoE) | 200.1 | 222.0 | 179.8 |
NVFP4 won once, on the Qwen MoE, through a fused-MoE CUTLASS kernel. On Gemma-4-31B it was slower than FP8. The architecture mattered far more than either format: MoE models decoded at 152 to 226 tokens per second, dense ones at 22 to 69.
The steel-man for NVFP4 is batch size, and this benchmark does not test it. AWQ is weight-only: its matrix multiplies still run in 16 bits. NVFP4 quantizes activations too, so at high concurrency, where serving becomes compute-bound, Blackwell's FP4 Tensor Cores could reverse the ranking. The benchmark ran one stream at a time, so it cannot settle that. Measure it on your concurrency before you pay the B200 rate for it.
PTQ Works for FP8 and Big Models. Small Models Need QAD.
Post-training quantization (PTQ) rounds a finished model's weights using a small calibration set, with no retraining. It is safe for FP8 and for large models, and it is where small models break. The 1B-to-405B study found FP8 the most robust method across tasks, found 70B-class models stable at 4-bit, and found smaller models suffered severe accuracy drops at 4-bit, and found quantized models often struggled with instruction-following and hallucination detection.
Quantization-aware training (QAT) fine-tunes the model with the rounding in the loop. NVIDIA's variant, quantization-aware distillation (QAD), trains the NVFP4 student to match a frozen BF16 teacher's output distribution. NVIDIA's technical report says QAD recovered near-BF16 accuracy on Nemotron 3 Nano, Nemotron Nano V2 and Llama Nemotron Super v1, and that it works without the full training data. It also argues plain QAT is unstable on models that went through SFT, RL and merging, which describes most models you would deploy.
Our rule of thumb: for a model under roughly 10B parameters going to 4-bit, prefer a vendor QAD or QAT checkpoint (NVIDIA ships a QAD NVFP4 checkpoint for its 30B MoE Nemotron 3.5 Lightning) over a community PTQ upload. You can browse what exists for your model on Hugging Face.
KV-Cache Quantization Is a Second Decision
KV-cache quantization compresses the attention memory that grows with context length, and it has its own quality cost that weight benchmarks do not isolate. In vLLM it is a separate flag, --kv-cache-dtype fp8. Without calibration every scale defaults to 1.0, and vLLM notes that sliding-window attention layers are more sensitive and can be excluded with --kv-cache-dtype-skip-layers.
Two findings should make you keep it apart from the weight decision. The independent benchmark's NVFP4 arms shipped with an FP8 KV cache, so their deltas mix both changes. And a long-context study accepted at EMNLP 2025 found 8-bit methods lost about 0.8% on average, while 4-bit methods lost up to 59% on long-context tasks, with worse results outside English. Short-prompt leaderboards will not show you that regression. If your workload is long documents, turn on one change at a time and evaluate each at your real context length.
For the memory side of the same trade-off, our KV-cache eviction analysis covers what to protect when the cache fills.
Who Should Not Pick Each Format
Skip FP8 if your fleet is A100s. Ampere has no FP8 Tensor Cores, so you get the memory saving through Marlin without the compute speedup, and INT4-AWQ will fit more on the same card. Skip it also if the model only fits at 4-bit.
Skip INT4-AWQ if your model is under about 10B parameters and its job is instruction-following. The scale study found severe 4-bit drops in smaller models and named instruction-following as a weak spot of quantized models. Use FP8 or a QAD checkpoint instead.
Skip NVFP4 if any part of your fleet is Hopper, if the model is dense, or if the only checkpoint available is a community PTQ upload. Skip it as a reason to buy Blackwell: on a dense model in the measured setup it was slower and bigger than AWQ, and B200 hours list at 68% over H100 on Lambda.
Skip GGUF for multi-user serving on datacenter GPUs. vLLM can load it, but llama.cpp is its home, and the GPU-native formats above have the throughput work behind them.
The loser is NVFP4 on dense models. It has the newest silicon behind it and the best spec sheet, and in the one reproducible test it decoded slower than AWQ on both dense models and shipped a file 60% larger on Gemma-4-31B. That can change as kernels mature, so rerun the benchmark when you upgrade vLLM.
How to Decide Without Regret
The choices people regret come from three places: a GPU generation locked in by a format, a capability that regressed where nobody measured, and a bit width taken at face value when sizing memory. These criteria predict them:
- Which GPUs will serve this model for the next three years? If Hopper is in the mix, NVFP4 is out.
- Does the model fit in FP8 with your KV cache at real context? If yes, stop there.
- Dense or MoE? Dense favours INT4-AWQ today. MoE on Blackwell is where NVFP4 competes.
- How big is the model? Under about 10B parameters, demand a QAD or QAT checkpoint for 4-bit.
- What is your concurrency? Single-stream results favour AWQ. High-batch serving on Blackwell is untested and may favour NVFP4.
What changes the answer: a vLLM release with faster NVFP4 dense kernels, an official QAD checkpoint for your model, or a high-batch benchmark that you run yourself.
This Week: Pull your 200 hardest production prompts, the ones that exercise the exact capability you deployed the model for, and score them on the BF16 model. That number is your baseline. Vendors publish throughput on their newest silicon and accuracy on aggregate leaderboards, and neither is your workload.
This Month: Run the same set against FP8 and one 4-bit checkpoint, changing weights and KV cache separately. Record decode tokens per second at your real concurrency and the checkpoint's actual size on disk.
Before the GPU Order: Put the measured tokens per second next to the per-GPU-hour rate for each generation you are considering. If NVFP4 does not beat INT4-AWQ on your model by more than the price gap, buy for FP8 and AWQ.
The Bottom Line
Treat a 4-bit throughput multiplier the way you would treat a vendor's benchmark on any other component: as a claim measured on their hardware, their model and their batch size. The format and the inference runtime get chosen together, since kernel support decides the speed. FP8 is close to free and runs natively on every NVIDIA generation since Ada. INT4-AWQ is the 4-bit default on any recent NVIDIA card. NVFP4 is a Blackwell MoE format whose case for dense models rests on kernels that have not shipped yet.
Run your 200 prompts through the quantized checkpoint before you sign the GPU order.
Continue Reading
- vLLM vs TensorRT-LLM vs SGLang: Default to vLLM
- Tencent Says 49B Active. Your Node Loads All 770B.
- Random Eviction Matched the Scorers. Protect the Prompt.
- AMD's Inference Discount Depends on a GPU You Can't Rent
- Best Air-Gapped LLM Stack: Apache Weights on vLLM Beat NVIDIA's Fee
- 3,400 Tokens/s Was Batch 1. At 100K Context, It's Batch 12.
