MLModels

Train Your Own LLM From Scratch

by angelos-p (GitHub)

BeginnerTutorialFreeWorkshop-length, six parts; the default ~10M-parameter model trains in ~45 min on an M3 Pro laptop

Build and train a ~10M-parameter GPT on your laptop in six guided parts, from tokenizer to text generation

Start LearningAdded Sep 28, 2026 · Updated Sep 28, 2026

Overview

Train Your Own LLM From Scratch is an open-source workshop by GitHub user angelos-p that simplifies Andrej Karpathy's nanoGPT into a guided, six-part build. The goal is a ~10-million-parameter GPT that trains on an ordinary laptop in under an hour. Part 1 covers tokenization: character-level encoding, building the vocabulary, and why Byte Pair Encoding performs poorly on small datasets. Part 2 implements the transformer itself: embeddings, self-attention, layer normalization and MLP blocks. Part 3 builds the training loop with the loss function, the AdamW optimizer, gradient clipping and learning-rate scheduling. Part 4 covers inference, including autoregressive decoding, temperature and top-k sampling. Part 5 trains on real data and has learners read loss curves and explore scaling behavior. Part 6 is a competition: learners train optimized models on alternative datasets to produce the best poetry. The stack is PyTorch, numpy, tqdm and tiktoken (for BPE on larger datasets), managed with uv, and it requires Python 3.12+. It runs on Apple Silicon (MPS), NVIDIA CUDA, CPU or Google Colab. The model comes in three sizes: Tiny (~0.5M parameters, ~5 min), Small (~4M, ~20 min) and Medium (~10M, ~45 min), with times measured on an M3 Pro. Learners fill in template files (model.py, train.py, generate.py) guided by six docs. The repo reached about 3,400 GitHub stars and hit the Hacker News front page in May 2026 with 478 points. Commenters there praised the documentation and noted that Raschka's LLMs-from-scratch and Stanford CS336 are the deeper next steps.

At a Glance

Topic
ML
Level
Beginner
Format
Tutorial
Cost
Free
Duration
Workshop-length, six parts; the default ~10M-parameter model trains in ~45 min on an M3 Pro laptop
Provider
angelos-p (GitHub)
Hands-on
Yes — code/exercises
Certificate
None

What You’ll Learn

  • ✓Build a character-level tokenizer and see why BPE struggles on small datasets
  • ✓Implement embeddings, self-attention, layer norm and MLP blocks of a GPT-style transformer in PyTorch
  • ✓Write a complete training loop with AdamW, gradient clipping and learning-rate scheduling
  • ✓Generate text with autoregressive decoding, temperature control and top-k sampling
  • ✓Train on real data and interpret loss curves to diagnose model behavior
  • ✓Run scaling experiments across ~0.5M, ~4M and ~10M parameter model configurations
  • ✓Optimize a small model on a new dataset in a poetry-generation competition

Highlights

  • •Trains end to end on a laptop (Apple MPS, CUDA, CPU or Colab); no cloud GPU budget needed
  • •Template files plus six step-by-step docs let you write the code yourself instead of reading a finished repo
  • •About 3,400 GitHub stars and a 478-point Hacker News front-page thread (May 2026)
  • •A strong first step before heavier from-scratch resources such as Stanford CS336 or Raschka's book

Who It’s For

Best For

  • ✓Developers who use LLM APIs daily and want to see how a model is actually trained
  • ✓Workshop organizers and meetup leads looking for a ready-made hands-on session
  • ✓Beginners who find nanoGPT or CS336 too steep as a starting point

Prerequisites

  • •Basic ability to read and run Python code
  • •A laptop or desktop with Python 3.12+ (or a Google Colab account)
  • •No prior machine learning background required

FAQ

What is Train Your Own LLM From Scratch?

Train Your Own LLM From Scratch is a free, hands-on workshop repository for developers who have never trained a language model. You write a GPT-style transformer in PyTorch and train it on Shakespeare on your own laptop, then use the working model to generate text and run scaling experiments.

Is Train Your Own LLM From Scratch free?

Train Your Own LLM From Scratch is free to access.

What level is Train Your Own LLM From Scratch for?

Train Your Own LLM From Scratch is aimed at a beginner audience. Recommended background: Basic ability to read and run Python code, A laptop or desktop with Python 3.12+ (or a Google Colab account), No prior machine learning background required.

How long does Train Your Own LLM From Scratch take?

Expect roughly Workshop-length, six parts; the default ~10M-parameter model trains in ~45 min on an M3 Pro laptop. Most learners work through it at their own pace.

What will I learn from Train Your Own LLM From Scratch?

You'll learn: Build a character-level tokenizer and see why BPE struggles on small datasets; Implement embeddings, self-attention, layer norm and MLP blocks of a GPT-style transformer in PyTorch; Write a complete training loop with AdamW, gradient clipping and learning-rate scheduling; Generate text with autoregressive decoding, temperature control and top-k sampling; Train on real data and interpret loss curves to diagnose model behavior; Run scaling experiments across ~0.5M, ~4M and ~10M parameter model configurations; Optimize a small model on a new dataset in a poetry-generation competition.

Topics

LLM from scratchGPTPyTorchnanoGPTtransformer trainingworkshop

Sources

This page was written from 2 sources, 1 on domains other than github.com.

  1. 1.github.com — llm from scratchvendor
  2. 2.news.ycombinator.com — item