ARENA: Alignment Research Engineer Accelerator Curriculum
by ARENA (London Initiative for Safe AI)
Free, code-first exercises in mechanistic interpretability, RL for LLMs, Inspect evals and alignment science, from transformers you build yourself to persona vectors
Overview
ARENA (Alignment Research Engineer Accelerator) is a programme of the London Initiative for Safe AI, a registered charity, and its full teaching material is free online at learn.arena.education and on GitHub (ARENA-education/ARENA_materials, about 1.3k stars and 873 forks). The curriculum has five chapters. Chapter 0, Fundamentals (7 sections), covers PyTorch, CNNs and ResNets, optimization, backpropagation, GANs and VAEs. Chapter 1, Transformer Interpretability (13 sections), starts with a transformer built from scratch and an intro to mechanistic interpretability with TransformerLens, then covers linear probes, function vectors and model steering, interpretability with sparse autoencoders, activation oracles, indirect object identification, SAE circuits, grokking, OthelloGPT and toy models of superposition. Chapter 2, Reinforcement Learning (6 sections), runs from tabular RL through DQN, policy gradients and PPO to RLHF on transformers and MCTS/AlphaZero. Chapter 3, LLM Evaluations (5 sections), covers eval design, model-generated datasets, running evals with the inspect-ai library, LLM agents and AI control. Chapter 4, Alignment Science (5 sections), was added in 2026 and replicates recent research: emergent misalignment, shutdown resistance and alignment faking case studies, reasoning-model interpretability via the Thought Anchors paper, Anthropic's Assistant Axis and persona vectors, and investigator agents built as a lightweight Petri on inspect-ai. The persona-vectors set loads Gemma and Qwen models and needs a CUDA GPU plus an OpenRouter key. ARENA 9.0 runs in London from October 5 to November 6, 2026, using the same material.
At a Glance
- Topic
- ML
- Level
- Advanced
- Format
- Course
- Cost
- Free
- Duration
- Self-paced; taught in person as a 4-5 week full-time bootcamp, and each 2026 alignment-science exercise set is about 1-2 days of work
- Provider
- ARENA (London Initiative for Safe AI)
- Hands-on
- Yes — code/exercises
- Certificate
- None
What You’ll Learn
- ✓Build and train a GPT-style transformer from scratch in PyTorch, then inspect it with TransformerLens
- ✓Find circuits such as induction heads and indirect object identification using activation patching
- ✓Train and interpret sparse autoencoders, and trace SAE circuits and attribution graphs in real models
- ✓Map persona space in model activations and extract trait-specific persona vectors with contrastive prompts
- ✓Implement DQN, PPO and RLHF on a transformer to see how LLM post-training works underneath
- ✓Design benchmarks, generate eval datasets with models, and run them with the inspect-ai library
- ✓Replicate emergent misalignment and alignment faking experiments on small open-weight models
- ✓Build investigator agents that probe target LLMs over multi-turn conversations to surface hidden behaviors
Highlights
- •Each exercise replicates a named paper (Geometry of Truth, Thought Anchors, Assistant Axis, Persona Vectors) in runnable code, so you finish with working research code
- •The 2026 alignment-science chapter adds eight new exercise sets, announced on LessWrong, that most interpretability courses do not cover yet
- •Used as the curriculum for ARENA's free in-person London bootcamps, so the notebooks are tested by each cohort and fixed through GitHub issues and PRs
- •Every set ships Colab notebooks with full solutions, so you can check your implementation against a reference
Who It’s For
Best For
- ✓ML engineers moving into interpretability, evals or alignment research roles
- ✓LLM engineers who want to understand post-training (RLHF, PPO) and model internals from code
- ✓People preparing an application to ARENA, MATS or a similar AI safety research programme
Prerequisites
- •Python proficiency and working familiarity with PyTorch
- •Linear algebra, calculus and probability at undergraduate level
- •A CUDA GPU (or Colab GPU) for the later chapters, plus an OpenRouter API key for some alignment-science sets
FAQ
What is ARENA: Alignment Research Engineer Accelerator Curriculum?
ARENA's open curriculum is a set of Colab and local Python exercise notebooks for ML engineers who want to work on how large language models function inside and how they misbehave. It suits people with solid PyTorch and linear algebra. Afterwards you can build and probe a transformer, train sparse autoencoders, run Inspect evals and replicate published alignment experiments.
Is ARENA: Alignment Research Engineer Accelerator Curriculum free?
ARENA: Alignment Research Engineer Accelerator Curriculum is free to access.
What level is ARENA: Alignment Research Engineer Accelerator Curriculum for?
ARENA: Alignment Research Engineer Accelerator Curriculum is aimed at a advanced audience. Recommended background: Python proficiency and working familiarity with PyTorch, Linear algebra, calculus and probability at undergraduate level, A CUDA GPU (or Colab GPU) for the later chapters, plus an OpenRouter API key for some alignment-science sets.
How long does ARENA: Alignment Research Engineer Accelerator Curriculum take?
Expect roughly Self-paced; taught in person as a 4-5 week full-time bootcamp, and each 2026 alignment-science exercise set is about 1-2 days of work. Most learners work through it at their own pace.
What will I learn from ARENA: Alignment Research Engineer Accelerator Curriculum?
You'll learn: Build and train a GPT-style transformer from scratch in PyTorch, then inspect it with TransformerLens; Find circuits such as induction heads and indirect object identification using activation patching; Train and interpret sparse autoencoders, and trace SAE circuits and attribution graphs in real models; Map persona space in model activations and extract trait-specific persona vectors with contrastive prompts; Implement DQN, PPO and RLHF on a transformer to see how LLM post-training works underneath; Design benchmarks, generate eval datasets with models, and run them with the inspect-ai library; Replicate emergent misalignment and alignment faking experiments on small open-weight models; Build investigator agents that probe target LLMs over multi-turn conversations to surface hidden behaviors.
Topics
Sources
This page was written from 6 sources, 4 on domains other than learn.arena.education.