ModelsFine-TuningML

NVIDIA Model Optimizer Documentation — Quantization, Pruning, Distillation and Speculative Decoding

by NVIDIA

IntermediateDocumentationFreeSelf-paced; ~1 hour per quick start, ~4-6 hours to work through the quantization and QAT guides

Compress LLMs to FP8, NVFP4 or INT4 and export them for TensorRT-LLM, vLLM or SGLang with NVIDIA's own optimization toolkit.

Start LearningAdded Oct 7, 2026 · Updated Oct 7, 2026

Overview

NVIDIA Model Optimizer, rebranded from TensorRT Model Optimizer in December 2025, ships its documentation in four parts. Getting Started covers installation on Linux (x86_64 and aarch64, Python 3.10 to 3.14, CUDA 12.x or 13.x, PyTorch 2.8 or later, ideally inside the TensorRT-LLM Docker image) and Windows, followed by quick starts for PTQ in PyTorch, ONNX and PyTorch-to-ONNX, PTQ on Windows, QAT and QAD, pruning, distillation, speculative decoding and sparsity. The Guides section holds the support matrix, a long quantization guide (basic concepts, best practices for choosing a method, PyTorch quantization, the quant_cfg configuration system, quantizing a custom Hugging Face model for TensorRT-LLM, compressing quantized models, and ONNX quantization on Linux and Windows), plus guides to quantization-aware training and distillation, saving and restoring, pruning, distillation, speculative decoding, sparsity, NAS, ONNX AutoCast and Autotune, recipes and the config system. Deployment pages cover TensorRT-LLM, ONNX Runtime and the unified Hugging Face checkpoint, which the GitHub README says also feeds vLLM and SGLang. Supported formats include FP8, NVFP4, INT8 and INT4 with AWQ, and the speculative decoding docs list EAGLE, Medusa, DFlash, DSpark and Domino. The project is actively maintained: in 2026 the repo added AutoQuantize for mixed-precision assignment (August), a blog on local-Hessian weight scales for NVFP4 accuracy and an end-to-end NVFP4 plus QAD tutorial for Qwen3.6-35B-A3B (September).

At a Glance

Topic
Models
Level
Intermediate
Format
Documentation
Cost
Free
Duration
Self-paced; ~1 hour per quick start, ~4-6 hours to work through the quantization and QAT guides
Provider
NVIDIA
Hands-on
Yes — code/exercises
Certificate
None

What You’ll Learn

  • ✓Install Model Optimizer with pip or inside the TensorRT-LLM Docker image and precompile its kernels
  • ✓Apply post-training quantization to PyTorch and ONNX models with calibration data
  • ✓Pick between FP8, NVFP4, INT8 and INT4 AWQ using the best-practice guide on accuracy trade-offs
  • ✓Write quant_cfg configurations to control which layers are quantized and at what precision
  • ✓Recover accuracy lost to low precision with quantization-aware training and quantization-aware distillation
  • ✓Prune and distill a large model into a smaller one using the dedicated quick starts
  • ✓Train speculative decoding draft heads such as EAGLE and Medusa to cut generation latency
  • ✓Export unified Hugging Face checkpoints for deployment on TensorRT-LLM, vLLM or SGLang

Highlights

  • •The vendor's own toolkit for NVFP4, the 4-bit format on recent NVIDIA GPUs, with 2026 tutorials on keeping NVFP4 accuracy high
  • •One library covers quantization, pruning, distillation, sparsity, NAS and speculative decoding, so you do not stitch together five projects
  • •About 5.2k GitHub stars, with runnable examples for Hugging Face PTQ, LLM QAT, pruning, speculative decoding, Megatron Bridge and diffusers
  • •Exports to the three serving engines most teams use (TensorRT-LLM, vLLM, SGLang) rather than locking you to one

Who It’s For

Best For

  • ✓Inference engineers cutting GPU memory and latency for self-hosted LLMs
  • ✓Teams deploying on NVIDIA Hopper or Blackwell hardware who want FP8 or NVFP4 checkpoints
  • ✓ML engineers who need to recover accuracy after quantization with QAT or distillation
  • ✓Developers adding speculative decoding to an existing serving stack

Prerequisites

  • •Python and PyTorch experience, including loading Hugging Face models
  • •Basic understanding of numeric precision and why quantization affects accuracy
  • •A Linux machine with an NVIDIA GPU and CUDA 12.x or 13.x (Windows is supported for some flows)
  • •Familiarity with at least one serving engine such as TensorRT-LLM, vLLM or SGLang

FAQ

What is NVIDIA Model Optimizer Documentation — Quantization, Pruning, Distillation and Speculative Decoding?

The official documentation for NVIDIA Model Optimizer (formerly TensorRT Model Optimizer), an Apache-2.0 library for shrinking and speeding up models before deployment. It is for engineers serving LLMs or diffusion models who need to apply post-training quantization, quantization-aware training, pruning, distillation, sparsity or speculative decoding and then export to an inference engine.

Is NVIDIA Model Optimizer Documentation — Quantization, Pruning, Distillation and Speculative Decoding free?

NVIDIA Model Optimizer Documentation — Quantization, Pruning, Distillation and Speculative Decoding is free to access.

What level is NVIDIA Model Optimizer Documentation — Quantization, Pruning, Distillation and Speculative Decoding for?

NVIDIA Model Optimizer Documentation — Quantization, Pruning, Distillation and Speculative Decoding is aimed at a intermediate audience. Recommended background: Python and PyTorch experience, including loading Hugging Face models, Basic understanding of numeric precision and why quantization affects accuracy, A Linux machine with an NVIDIA GPU and CUDA 12.x or 13.x (Windows is supported for some flows), Familiarity with at least one serving engine such as TensorRT-LLM, vLLM or SGLang.

How long does NVIDIA Model Optimizer Documentation — Quantization, Pruning, Distillation and Speculative Decoding take?

Expect roughly Self-paced; ~1 hour per quick start, ~4-6 hours to work through the quantization and QAT guides. Most learners work through it at their own pace.

What will I learn from NVIDIA Model Optimizer Documentation — Quantization, Pruning, Distillation and Speculative Decoding?

You'll learn: Install Model Optimizer with pip or inside the TensorRT-LLM Docker image and precompile its kernels; Apply post-training quantization to PyTorch and ONNX models with calibration data; Pick between FP8, NVFP4, INT8 and INT4 AWQ using the best-practice guide on accuracy trade-offs; Write quant_cfg configurations to control which layers are quantized and at what precision; Recover accuracy lost to low precision with quantization-aware training and quantization-aware distillation; Prune and distill a large model into a smaller one using the dedicated quick starts; Train speculative decoding draft heads such as EAGLE and Medusa to cut generation latency; Export unified Hugging Face checkpoints for deployment on TensorRT-LLM, vLLM or SGLang.

Topics

quantizationnvfp4model-compressionspeculative-decodingtensorrt-llmdistillation

Sources

This page was written from 4 sources, 1 on domains other than nvidia.github.io.

  1. 1.nvidia.github.io — Model Optimizervendor
  2. 2.nvidia.github.io — 1 quantizationvendor
  3. 3.nvidia.github.io — installation for Linuxvendor
  4. 4.github.com — Model Optimizer