MLModelsFine-Tuning

The Ultra-Scale Playbook: Training LLMs on GPU Clusters

by Hugging Face (nanotron team)

AdvancedGuideFreeThe playbook's own estimated reading time is 2-4 days; longer if you work through the picotron code

The free, book-length account of how LLM pretraining actually runs on 512 GPUs — built on 4,000+ experiments.

Start LearningAdded Jul 4, 2026 · Updated Aug 21, 2026

Overview

Hugging Face's nanotron team published this in March 2025 and it remains the most complete free account of how large-model pretraining actually runs on a cluster. It is grounded in over 4,000 scaling experiments on up to 512 GPUs — more than 16,000 runs including tests — and its own stated reading time is two to four days, which is an honest signal that this is a book presented as a web page rather than a blog post. It opens on a single GPU with memory accounting: weights, gradients, optimizer states and activations, then activation recomputation and gradient accumulation as the first levers. From there it introduces the parallelism axes one at a time, each with the communication pattern it adds and the point at which it stops paying for itself: data parallelism, ZeRO stages 1 through 3, tensor parallelism, sequence parallelism, context parallelism, pipeline parallelism with its bubble-scheduling variants, and expert parallelism for mixture-of-experts models — together the 5D parallelism the playbook is known for. Later material covers mixed precision, fused and fast CUDA kernels, flash attention, and overlapping computation with communication to keep GPUs from idling. Interactive widgets let you compute a transformer's memory breakdown in the page. Two repositories accompany it: picotron, self-contained educational implementations, and nanotron, the production trainer Hugging Face uses internally.

At a Glance

Topic
ML
Level
Advanced
Format
Guide
Cost
Free
Duration
The playbook's own estimated reading time is 2-4 days; longer if you work through the picotron code
Provider
Hugging Face (nanotron team)
Hands-on
Yes — code/exercises
Certificate
None

What You’ll Learn

  • ✓Account for GPU memory precisely across weights, gradients, optimizer states and activations
  • ✓Apply activation recomputation and gradient accumulation before reaching for more GPUs
  • ✓Understand what each ZeRO stage shards and what communication cost that sharding buys
  • ✓Choose between tensor, sequence, context, pipeline and expert parallelism for a given model shape
  • ✓Recognise pipeline bubbles and the scheduling variants that shrink them at scale
  • ✓Overlap communication with computation so interconnect latency stops idling the accelerators
  • ✓Read picotron's self-contained implementations to see each technique in runnable code

Highlights

  • •Backed by 4,000+ real scaling experiments on up to 512 GPUs rather than theory alone
  • •Interactive widgets compute a transformer's memory breakdown directly in the page
  • •Ships with picotron for learning and nanotron, the trainer Hugging Face actually uses in production
  • •Explains where each parallelism strategy stops paying off, not just how to enable it
  • •Completely free and open, covering material otherwise scattered across frontier-lab papers

Who It’s For

Best For

  • ✓ML engineers running or planning multi-node pretraining and large fine-tuning jobs
  • ✓Infrastructure engineers sizing GPU clusters and interconnect for training workloads
  • ✓Researchers who want to understand 5D parallelism at implementation depth

Prerequisites

  • •Solid PyTorch and a working understanding of transformer training loops
  • •Familiarity with GPU memory concepts and at least single-node distributed training
  • •Access to multiple GPUs if you intend to reproduce any of the experiments

FAQ

What is The Ultra-Scale Playbook: Training LLMs on GPU Clusters?

Hugging Face's open guide to distributed LLM training, built on more than 4,000 scaling experiments on up to 512 GPUs. It starts with memory accounting on a single GPU and builds up the parallelism axes one at a time — data, ZeRO, tensor, sequence, context, pipeline and expert — explaining the communication cost each one introduces and where it stops paying off, with interactive widgets and two companion code repositories.

Is The Ultra-Scale Playbook: Training LLMs on GPU Clusters free?

The Ultra-Scale Playbook: Training LLMs on GPU Clusters is free to access.

What level is The Ultra-Scale Playbook: Training LLMs on GPU Clusters for?

The Ultra-Scale Playbook: Training LLMs on GPU Clusters is aimed at a advanced audience. Recommended background: Solid PyTorch and a working understanding of transformer training loops, Familiarity with GPU memory concepts and at least single-node distributed training, Access to multiple GPUs if you intend to reproduce any of the experiments.

How long does The Ultra-Scale Playbook: Training LLMs on GPU Clusters take?

Expect roughly The playbook's own estimated reading time is 2-4 days; longer if you work through the picotron code. Most learners work through it at their own pace.

What will I learn from The Ultra-Scale Playbook: Training LLMs on GPU Clusters?

You'll learn: Account for GPU memory precisely across weights, gradients, optimizer states and activations; Apply activation recomputation and gradient accumulation before reaching for more GPUs; Understand what each ZeRO stage shards and what communication cost that sharding buys; Choose between tensor, sequence, context, pipeline and expert parallelism for a given model shape; Recognise pipeline bubbles and the scheduling variants that shrink them at scale; Overlap communication with computation so interconnect latency stops idling the accelerators; Read picotron's self-contained implementations to see each technique in runnable code.

Topics

distributed-training5d-parallelismzeropretraininggpu-clustersnanotron

Sources

This page was written from 3 sources, 2 on domains other than huggingface.co.

  1. 1.huggingface.co — ultrascale playbookvendor
  2. 2.infoq.com — huggingface ultra scale playbook
  3. 3.github.com — nanotron