Train Your Own LLM From Scratch
by angelos-p (GitHub)
Build and train a ~10M-parameter GPT on your laptop in six guided parts, from tokenizer to text generation
Overview
Train Your Own LLM From Scratch is an open-source workshop by GitHub user angelos-p that simplifies Andrej Karpathy's nanoGPT into a guided, six-part build. The goal is a ~10-million-parameter GPT that trains on an ordinary laptop in under an hour. Part 1 covers tokenization: character-level encoding, building the vocabulary, and why Byte Pair Encoding performs poorly on small datasets. Part 2 implements the transformer itself: embeddings, self-attention, layer normalization and MLP blocks. Part 3 builds the training loop with the loss function, the AdamW optimizer, gradient clipping and learning-rate scheduling. Part 4 covers inference, including autoregressive decoding, temperature and top-k sampling. Part 5 trains on real data and has learners read loss curves and explore scaling behavior. Part 6 is a competition: learners train optimized models on alternative datasets to produce the best poetry. The stack is PyTorch, numpy, tqdm and tiktoken (for BPE on larger datasets), managed with uv, and it requires Python 3.12+. It runs on Apple Silicon (MPS), NVIDIA CUDA, CPU or Google Colab. The model comes in three sizes: Tiny (~0.5M parameters, ~5 min), Small (~4M, ~20 min) and Medium (~10M, ~45 min), with times measured on an M3 Pro. Learners fill in template files (model.py, train.py, generate.py) guided by six docs. The repo reached about 3,400 GitHub stars and hit the Hacker News front page in May 2026 with 478 points. Commenters there praised the documentation and noted that Raschka's LLMs-from-scratch and Stanford CS336 are the deeper next steps.
At a Glance
- Topic
- ML
- Level
- Beginner
- Format
- Tutorial
- Cost
- Free
- Duration
- Workshop-length, six parts; the default ~10M-parameter model trains in ~45 min on an M3 Pro laptop
- Provider
- angelos-p (GitHub)
- Hands-on
- Yes — code/exercises
- Certificate
- None
What You’ll Learn
- ✓Build a character-level tokenizer and see why BPE struggles on small datasets
- ✓Implement embeddings, self-attention, layer norm and MLP blocks of a GPT-style transformer in PyTorch
- ✓Write a complete training loop with AdamW, gradient clipping and learning-rate scheduling
- ✓Generate text with autoregressive decoding, temperature control and top-k sampling
- ✓Train on real data and interpret loss curves to diagnose model behavior
- ✓Run scaling experiments across ~0.5M, ~4M and ~10M parameter model configurations
- ✓Optimize a small model on a new dataset in a poetry-generation competition
Highlights
- •Trains end to end on a laptop (Apple MPS, CUDA, CPU or Colab); no cloud GPU budget needed
- •Template files plus six step-by-step docs let you write the code yourself instead of reading a finished repo
- •About 3,400 GitHub stars and a 478-point Hacker News front-page thread (May 2026)
- •A strong first step before heavier from-scratch resources such as Stanford CS336 or Raschka's book
Who It’s For
Best For
- ✓Developers who use LLM APIs daily and want to see how a model is actually trained
- ✓Workshop organizers and meetup leads looking for a ready-made hands-on session
- ✓Beginners who find nanoGPT or CS336 too steep as a starting point
Prerequisites
- •Basic ability to read and run Python code
- •A laptop or desktop with Python 3.12+ (or a Google Colab account)
- •No prior machine learning background required
FAQ
What is Train Your Own LLM From Scratch?
Train Your Own LLM From Scratch is a free, hands-on workshop repository for developers who have never trained a language model. You write a GPT-style transformer in PyTorch and train it on Shakespeare on your own laptop, then use the working model to generate text and run scaling experiments.
Is Train Your Own LLM From Scratch free?
Train Your Own LLM From Scratch is free to access.
What level is Train Your Own LLM From Scratch for?
Train Your Own LLM From Scratch is aimed at a beginner audience. Recommended background: Basic ability to read and run Python code, A laptop or desktop with Python 3.12+ (or a Google Colab account), No prior machine learning background required.
How long does Train Your Own LLM From Scratch take?
Expect roughly Workshop-length, six parts; the default ~10M-parameter model trains in ~45 min on an M3 Pro laptop. Most learners work through it at their own pace.
What will I learn from Train Your Own LLM From Scratch?
You'll learn: Build a character-level tokenizer and see why BPE struggles on small datasets; Implement embeddings, self-attention, layer norm and MLP blocks of a GPT-style transformer in PyTorch; Write a complete training loop with AdamW, gradient clipping and learning-rate scheduling; Generate text with autoregressive decoding, temperature control and top-k sampling; Train on real data and interpret loss curves to diagnose model behavior; Run scaling experiments across ~0.5M, ~4M and ~10M parameter model configurations; Optimize a small model on a new dataset in a poetry-generation competition.
Topics
Sources
This page was written from 2 sources, 1 on domains other than github.com.