Scaling Laws, Carefully
by Lilian Weng (Lil'Log)
Kaplan vs. Chinchilla, data-limited scaling and the practical traps of fitting scaling laws, in one careful review
Overview
Scaling Laws, Carefully is a June 24, 2026 post on Lil'Log by Lilian Weng, about a 25-minute read. It starts with early work on loss predictability: Amari et al. (1992) on power-law learning curves, Hestness et al. (2017) on predictable error decay across translation, vision, language modeling and speech, and Rosenfeld et al. (2020) on modeling error as a joint function of model and data size. The data-infinite section covers Kaplan et al. (2020), including N_opt ∝ C^0.73 and the C ≈ 6ND compute approximation, then the three Chinchilla methods from Hoffmann et al. (2022): fixed model sizes with varying token budgets, IsoFLOP profiles and a parametric fit with Huber loss and L-BFGS, which lead to N_opt ∝ C^0.5. It then explains how Pearce and Song (2024) reconcile the two through embedding parameters, and why power laws appear at all (Sharma and Kaplan 2020, Michaud et al. 2023). The data-limited section covers repeated data, from Hernandez et al. (2022) and Muennighoff et al. (2023) to an overfitting-penalty law from Lovelace et al. (2026). It ends with fitting pitfalls, including the Besiroglu et al. (2024) critique of Chinchilla's Method 3, and a toy simulation of how loss precision, noise and fit region change predictions.
At a Glance
- Topic
- Models
- Level
- Advanced
- Format
- Guide
- Cost
- Free
- Duration
- ~25 min read; longer if you work through the formulas and toy simulation
- Provider
- Lilian Weng (Lil'Log)
- Hands-on
- No
- Certificate
- None
What You’ll Learn
- ✓How early work by Hestness and Rosenfeld showed generalization error decays predictably with data and model size
- ✓What Kaplan et al. found about compute-optimal model size and the C ≈ 6ND training FLOPs approximation
- ✓How Chinchilla's three methods (fixed sizes, IsoFLOP profiles, parametric fit) reached equal scaling of model and data
- ✓Why Kaplan and Chinchilla disagree, including the role of embedding parameters in small models
- ✓The leading explanations for why loss follows a power law, from data manifolds to quantized skills
- ✓How data repetition changes scaling once unique high-quality tokens run out, including overfitting-penalty laws
- ✓Which fitting choices (loss precision, noise, fit region, optimizer termination) distort scaling-law extrapolations
Highlights
- •Gives the actual loss formulas for each law rather than paraphrasing them
- •Includes 2026 work on data-limited scaling, which older scaling-law explainers predate
- •Ends with a toy simulation showing how small fitting choices swing predictions, which few tutorials cover
- •Reached 88 points and 20 comments on Hacker News
- •Free, from the author of the widely cited Lil'Log ML surveys
Who It’s For
Best For
- ✓ML engineers planning pretraining or continued-pretraining compute budgets
- ✓Researchers who fit or report scaling laws and want to avoid known methodological traps
- ✓Engineers evaluating vendor or paper claims about compute-optimal models
- ✓Readers who know Chinchilla by name and want the derivations behind it
Prerequisites
- •Comfort with log-log plots, power laws and basic optimization
- •Familiarity with transformer language model training: parameters, tokens and FLOPs
- •Willingness to follow the cited papers for full derivations
FAQ
What is Scaling Laws, Carefully?
A technical review by Lilian Weng of neural scaling laws for language models, from early learning-curve work through Kaplan, Chinchilla and recent data-limited laws. It is for ML engineers and researchers who plan training runs or read scaling claims. Afterwards you can explain why Kaplan and Chinchilla disagree and spot fitting mistakes that distort a scaling-law prediction.
Is Scaling Laws, Carefully free?
Scaling Laws, Carefully is free to access.
What level is Scaling Laws, Carefully for?
Scaling Laws, Carefully is aimed at a advanced audience. Recommended background: Comfort with log-log plots, power laws and basic optimization, Familiarity with transformer language model training: parameters, tokens and FLOPs, Willingness to follow the cited papers for full derivations.
How long does Scaling Laws, Carefully take?
Expect roughly ~25 min read; longer if you work through the formulas and toy simulation. Most learners work through it at their own pace.
What will I learn from Scaling Laws, Carefully?
You'll learn: How early work by Hestness and Rosenfeld showed generalization error decays predictably with data and model size; What Kaplan et al. found about compute-optimal model size and the C ≈ 6ND training FLOPs approximation; How Chinchilla's three methods (fixed sizes, IsoFLOP profiles, parametric fit) reached equal scaling of model and data; Why Kaplan and Chinchilla disagree, including the role of embedding parameters in small models; The leading explanations for why loss follows a power law, from data manifolds to quantized skills; How data repetition changes scaling once unique high-quality tokens run out, including overfitting-penalty laws; Which fitting choices (loss precision, noise, fit region, optimizer termination) distort scaling-law extrapolations.
Topics
Sources
This page was written from 2 sources, 1 on domains other than lilianweng.github.io.