FP8 vs NVFP4 vs INT4-AWQ: AWQ Outran NVFP4 on 3 of 4 Models

Run FP8 by default. For 4-bit, INT4-AWQ beat NVFP4 on speed on three of four models in an independent Blackwell benchmark, and NVIDIA's own NVFP4 Gemma file was nearly the size of FP8.

By Rajesh Beri·October 9, 2026·12 min read
Share:
A server technician's gloved hand holding a GPU accelerator card above an open rack server, with two other GPU cards of different generations lying side by side on an anti-static mat next to it.

Illustration generated using AI

A vendor's 4-bit throughput multiplier hides two decisions: which GPU generation you are committing to, and which capability your model loses. For most enterprise inference, run FP8 on Hopper or Blackwell: in an independent 16-arm benchmark it moved MMLU-Pro by 0.7 points or less on every model tested. When you need 4-bit to fit a bigger model or a longer context, use INT4-AWQ on mature Marlin kernels: on a Blackwell workstation card it decoded faster than NVFP4 on three of four models and produced the smaller file on both dense ones. NVFP4 loses on dense models today. It earns its place on mixture-of-experts models on Blackwell, ideally from a checkpoint NVIDIA recovered with quantization-aware distillation. GGUF belongs on laptops, CPUs and Macs, not in a GPU serving fleet.

Format Runs natively on MMLU-Pro change vs BF16 (4 models) Decode, Qwen3.6-27B, 1 stream File size, Gemma-4-31B Pick it when Skip it when
FP8 (E4M3) Ada, Hopper, Blackwell, AMD GPU -0.1 to -0.7 48.8 tok/s 33.3 GB You have the memory. This is the default. You are on Ampere or older
INT4-AWQ Turing through Hopper in vLLM's table, plus Blackwell, Intel GPU, x86 CPU -0.1 to -1.5 69.0 tok/s 20.5 GB You need 4-bit on a dense model, on any recent NVIDIA card You serve at high batch on Blackwell and have measured NVFP4 winning there
NVFP4 Blackwell only +0.2 to -2.4 55.6 tok/s 32.7 GB MoE models on Blackwell, QAD checkpoints Your fleet includes Hopper, or the model is small and dense
GGUF (llama.cpp) CPU, Apple silicon, most GPUs Not measured in this benchmark Not measured Not measured Laptops, edge boxes, air-gapped desktops You run a multi-user GPU service

The workload behind the numbers: one RTX PRO 6000 Blackwell Max-Q (96 GB) running vLLM 0.22.1, single stream, 64k context, published June 8, 2026 by Stefano Schotten. Its code is MIT-licensed with both full runs committed under results/ and results-rerun/, and the checkpoints are pinned by SHA. That makes it the only cross-format comparison we found that you can rerun yourself.


What Each Format Is, and the GPU It Commits You To

Quantization is storing a model's weights, and sometimes its activations, in fewer bits than the 16 it was trained in. Each format is tied to specific silicon, so picking one is a hardware decision.

FP8 is an 8-bit floating-point format. NVIDIA added FP8 Tensor Cores with the H100, in two variants: E4M3 (4 exponent bits, 3 mantissa bits) for precision, and E5M2 for range. Inference uses E4M3. In vLLM's hardware table, full W8A8 FP8 is supported on Ada and Hopper among NVIDIA columns, plus AMD GPUs. Ampere only gets FP8 weights through the Marlin weight-only path.

NVFP4 is NVIDIA's 4-bit float for Blackwell. Each block of 16 values shares an FP8 E4M3 scale, with a second FP32 scale per tensor, and NVIDIA's fifth-generation Tensor Cores execute it in hardware. FP4 hardware support arrived with Blackwell, so an NVFP4 checkpoint commits you to Blackwell.

INT4-AWQ is 4-bit integer weights with 16-bit activations. AWQ finds the roughly 1% of weights that matter most from activation statistics and scales them up before rounding. GPTQ is the older alternative that compensates rounding error layer by layer. vLLM supports AWQ from Turing through Hopper and on Intel GPUs and x86 CPUs, which makes it the most portable 4-bit option. In a study of models from 1B to 405B parameters, AWQ tended to beat GPTQ on accuracy.

AutoRound is a recipe that emits these formats, and you should evaluate it as one. Intel's Apache-2.0 toolkit exports to AutoGPTQ, AutoAWQ, GGUF and llm-compressor, and supports NVFP4 and MXFP4 schemes.

GGUF is llama.cpp's file format, built for CPUs and consumer hardware. Its popular Q4_K_M preset mixes precisions and lands near 4.9 bits per weight on Llama-3.1-8B, above the "4-bit" label.

The hardware commitment has a price. On Lambda's on-demand list, checked October 9, 2026, an 8-GPU H100 SXM instance is $3.99 per GPU-hour and an 8-GPU B200 instance is $6.69 per GPU-hour, 68% more. A format that only runs natively on Blackwell has to pay for that gap in throughput you actually measure.


Where the Accuracy Loss Actually Lands

Four-bit costs little on average, and nearly all of that cost lands in knowledge recall. Averaged over MMLU-Pro, GSM8K, IFEval, HumanEval and MBPP, NVFP4 cost at most 0.6 points on the dense models in the independent benchmark. Math, code and instruction-following sat near their ceilings.

MMLU-Pro is where it shows. On Qwen3.6-27B, MMLU-Pro fell 0.4 points under FP8, 1.5 under INT4-AWQ and 2.4 under NVFP4. On Gemma-4-31B the drops were 0.7, 1.0 and 1.6. The two MoE models barely moved: Qwen3.6-35B-A3B lost 0.5 under NVFP4, and Gemma-4-26B-A4B gained 0.2. The author's reading is that NVFP4 and INT4-AWQ are a wash at equal bytes per parameter, and that the recipe decides the result more than the number format does.

NVIDIA's own numbers point the same way at larger scale. On DeepSeek-R1-0528, NVFP4 scored within one point of FP8 on MMLU-Pro (84% vs 85%), GPQA Diamond and LiveCodeBench, and two points higher on AIME 2024.

Two caveats change how you read this. The benchmark's retest found a mean drift of 0.35 points across 80 scores, with MMLU-Pro moving up to 1.7 points between identical runs, so sub-point gaps are inside the noise. And when the author checked NVIDIA's published deltas, they matched on Qwen3.6-35B-A3B (-0.57 measured vs -0.60 published) and diverged on Gemma-4-31B: -1.57 measured against NVIDIA's -0.31.


Why an NVFP4 File Can Be Bigger Than an AWQ File

NVFP4's scales cost real bits, and NVIDIA's checkpoints leave some layers at full precision, so "4-bit" files vary widely in size. NVIDIA puts NVFP4 at about 4.5 bits per value, roughly 1.8x smaller than FP8.

The files on disk tell a different story. In the benchmark's size table, Gemma-4-31B came to 20.5 GB as INT4-AWQ and 32.7 GB as NVIDIA's official NVFP4, within 0.6 GB of the 33.3 GB FP8 file. NVIDIA's recipe keeps most attention layers in BF16. Qwen3.6-27B was 21.9 GB as AWQ and 26.4 GB as the community NVFP4. On Gemma-4-26B-A4B, AWQ was again smaller (17.2 GB against 18.8 GB). NVFP4 produced the smallest file on one model of four, Qwen3.6-35B-A3B, at 23.5 GB against AWQ's 25.5 GB.

If you size GPU memory from the nominal bit width, you will be wrong by up to 12 GB on a 31B model. Size it from the checkpoint you will actually deploy.


Throughput Follows Kernel Maturity, Not the Bit Width

The format with the better on-paper spec lost most of the speed tests, because INT4-AWQ runs on the mature Marlin kernels. Single-stream decode on the Blackwell card, per the benchmark:

Model FP8 INT4-AWQ NVFP4
Qwen3.6-27B (dense) 48.8 69.0 55.6
Gemma-4-31B-it (dense) 42.0 63.9 40.8
Qwen3.6-35B-A3B (MoE) 221.4 194.2 226.3
Gemma-4-26B-A4B (MoE) 200.1 222.0 179.8

NVFP4 won once, on the Qwen MoE, through a fused-MoE CUTLASS kernel. On Gemma-4-31B it was slower than FP8. The architecture mattered far more than either format: MoE models decoded at 152 to 226 tokens per second, dense ones at 22 to 69.

The steel-man for NVFP4 is batch size, and this benchmark does not test it. AWQ is weight-only: its matrix multiplies still run in 16 bits. NVFP4 quantizes activations too, so at high concurrency, where serving becomes compute-bound, Blackwell's FP4 Tensor Cores could reverse the ranking. The benchmark ran one stream at a time, so it cannot settle that. Measure it on your concurrency before you pay the B200 rate for it.


PTQ Works for FP8 and Big Models. Small Models Need QAD.

Post-training quantization (PTQ) rounds a finished model's weights using a small calibration set, with no retraining. It is safe for FP8 and for large models, and it is where small models break. The 1B-to-405B study found FP8 the most robust method across tasks, found 70B-class models stable at 4-bit, and found smaller models suffered severe accuracy drops at 4-bit, and found quantized models often struggled with instruction-following and hallucination detection.

Quantization-aware training (QAT) fine-tunes the model with the rounding in the loop. NVIDIA's variant, quantization-aware distillation (QAD), trains the NVFP4 student to match a frozen BF16 teacher's output distribution. NVIDIA's technical report says QAD recovered near-BF16 accuracy on Nemotron 3 Nano, Nemotron Nano V2 and Llama Nemotron Super v1, and that it works without the full training data. It also argues plain QAT is unstable on models that went through SFT, RL and merging, which describes most models you would deploy.

Our rule of thumb: for a model under roughly 10B parameters going to 4-bit, prefer a vendor QAD or QAT checkpoint (NVIDIA ships a QAD NVFP4 checkpoint for its 30B MoE Nemotron 3.5 Lightning) over a community PTQ upload. You can browse what exists for your model on Hugging Face.


KV-Cache Quantization Is a Second Decision

KV-cache quantization compresses the attention memory that grows with context length, and it has its own quality cost that weight benchmarks do not isolate. In vLLM it is a separate flag, --kv-cache-dtype fp8. Without calibration every scale defaults to 1.0, and vLLM notes that sliding-window attention layers are more sensitive and can be excluded with --kv-cache-dtype-skip-layers.

Two findings should make you keep it apart from the weight decision. The independent benchmark's NVFP4 arms shipped with an FP8 KV cache, so their deltas mix both changes. And a long-context study accepted at EMNLP 2025 found 8-bit methods lost about 0.8% on average, while 4-bit methods lost up to 59% on long-context tasks, with worse results outside English. Short-prompt leaderboards will not show you that regression. If your workload is long documents, turn on one change at a time and evaluate each at your real context length.

For the memory side of the same trade-off, our KV-cache eviction analysis covers what to protect when the cache fills.


Who Should Not Pick Each Format

Skip FP8 if your fleet is A100s. Ampere has no FP8 Tensor Cores, so you get the memory saving through Marlin without the compute speedup, and INT4-AWQ will fit more on the same card. Skip it also if the model only fits at 4-bit.

Skip INT4-AWQ if your model is under about 10B parameters and its job is instruction-following. The scale study found severe 4-bit drops in smaller models and named instruction-following as a weak spot of quantized models. Use FP8 or a QAD checkpoint instead.

Skip NVFP4 if any part of your fleet is Hopper, if the model is dense, or if the only checkpoint available is a community PTQ upload. Skip it as a reason to buy Blackwell: on a dense model in the measured setup it was slower and bigger than AWQ, and B200 hours list at 68% over H100 on Lambda.

Skip GGUF for multi-user serving on datacenter GPUs. vLLM can load it, but llama.cpp is its home, and the GPU-native formats above have the throughput work behind them.

The loser is NVFP4 on dense models. It has the newest silicon behind it and the best spec sheet, and in the one reproducible test it decoded slower than AWQ on both dense models and shipped a file 60% larger on Gemma-4-31B. That can change as kernels mature, so rerun the benchmark when you upgrade vLLM.


How to Decide Without Regret

The choices people regret come from three places: a GPU generation locked in by a format, a capability that regressed where nobody measured, and a bit width taken at face value when sizing memory. These criteria predict them:

  1. Which GPUs will serve this model for the next three years? If Hopper is in the mix, NVFP4 is out.
  2. Does the model fit in FP8 with your KV cache at real context? If yes, stop there.
  3. Dense or MoE? Dense favours INT4-AWQ today. MoE on Blackwell is where NVFP4 competes.
  4. How big is the model? Under about 10B parameters, demand a QAD or QAT checkpoint for 4-bit.
  5. What is your concurrency? Single-stream results favour AWQ. High-batch serving on Blackwell is untested and may favour NVFP4.

What changes the answer: a vLLM release with faster NVFP4 dense kernels, an official QAD checkpoint for your model, or a high-batch benchmark that you run yourself.

This Week: Pull your 200 hardest production prompts, the ones that exercise the exact capability you deployed the model for, and score them on the BF16 model. That number is your baseline. Vendors publish throughput on their newest silicon and accuracy on aggregate leaderboards, and neither is your workload.

This Month: Run the same set against FP8 and one 4-bit checkpoint, changing weights and KV cache separately. Record decode tokens per second at your real concurrency and the checkpoint's actual size on disk.

Before the GPU Order: Put the measured tokens per second next to the per-GPU-hour rate for each generation you are considering. If NVFP4 does not beat INT4-AWQ on your model by more than the price gap, buy for FP8 and AWQ.


The Bottom Line

Treat a 4-bit throughput multiplier the way you would treat a vendor's benchmark on any other component: as a claim measured on their hardware, their model and their batch size. The format and the inference runtime get chosen together, since kernel support decides the speed. FP8 is close to free and runs natively on every NVIDIA generation since Ada. INT4-AWQ is the 4-bit default on any recent NVIDIA card. NVFP4 is a Blackwell MoE format whose case for dense models rests on kernels that have not shipped yet.

Run your 200 prompts through the quantized checkpoint before you sign the GPU order.

Continue Reading

Share:

Frequently Asked Questions

Is NVFP4 more accurate than INT4-AWQ?

Not measurably. In an independent, reproducible benchmark on four models, NVFP4 and INT4-AWQ landed within noise of each other on average, and both lost most of their accuracy on knowledge recall (MMLU-Pro) rather than math, code or instruction-following. The quantization recipe mattered more than the number format.

Does NVFP4 run on H100 GPUs?

NVFP4 is executed natively by the fifth-generation Tensor Cores in NVIDIA's Blackwell GPUs. Choosing NVFP4 checkpoints commits your serving fleet to Blackwell. On Hopper, FP8 or INT4-AWQ are the practical choices.

Why is my NVFP4 model file bigger than the AWQ version?

NVFP4 stores an FP8 scale for every 16 values, about 4.5 bits per value, and some official checkpoints keep attention layers in BF16. NVIDIA's NVFP4 Gemma-4-31B is 32.7 GB on disk against 20.5 GB for an INT4-AWQ build of the same model.

Should I quantize the KV cache at the same time as the weights?

No. Treat it as a separate change and evaluate it at your real context length. vLLM enables it with --kv-cache-dtype fp8, uncalibrated scales default to 1.0, and long-context studies show quality losses that short-prompt benchmarks miss.

Is post-training quantization safe for small models?

FP8 PTQ is generally safe. At 4-bit, small models can lose accuracy sharply, and quantized models often struggle with instruction-following. For models under roughly 10B parameters, prefer a checkpoint recovered with quantization-aware training or NVIDIA's quantization-aware distillation.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →