CS 2881R: AI Safety (Harvard, Fall 2026)
by Harvard University (Boaz Barak)
Harvard's graduate AI safety seminar, 2026 edition: RL post-training, interpretability, cyber and bio risk, taught with guests from Anthropic and OpenAI.
Overview
CS 2881R: AI Safety is a Harvard computer science course taught by Boaz Barak, a Harvard CS professor. The Fall 2026 offering is its second run, after the first in Fall 2025. It meets weekly from September 3 to December 3, 2026, for 13 in-person lectures of about 2 hours 45 minutes each. In each session a group of students presents an experiment, and the class discusses it. The public schedule runs from an introduction to cyber capabilities (guest Nicholas Carlini of Anthropic) and then modern LLM training and inference. Next come recursive self-improvement and AI trajectories (Dwarkesh Patel, Daniel Kokotajlo) and the economic impact of AI (Chad Jones, Erik Brynjolfsson). Later sessions cover reinforcement learning for post-training and alignment (John Schulman), model policies (Ziad Reslan), open-source models (Nathan Lambert), AI interpretability (Jack Lindsey), alignment in the age of recursive self-improvement (Jakub Pachocki), and AI biosecurity and threat modeling (Luca Righetti). Two slots are still TBD. Enrolled students give experiment presentations, write scribe notes and do a final project. The 2025 edition added a paper-replication midterm and a NeurIPS-style final paper. The course encourages heavy use of generative and agentic AI for coursework. The syllabus says lectures will be recorded and published online where technically possible, and Lecture 3 (Modern LLM Training and Inference) is already on YouTube. The 2025 syllabus, videos, rubrics and student papers are public too.
At a Glance
- Topic
- Models
- Level
- Advanced
- Format
- Course
- Cost
- Free
- Duration
- 13 weekly lectures (~2h45m each), Sep 3 – Dec 3, 2026; recordings posted to YouTube, sometimes with a delay
- Provider
- Harvard University (Boaz Barak)
- Hands-on
- Yes — code/exercises
- Certificate
- None
What You’ll Learn
- ✓How modern LLMs are pre-trained, post-trained and served, as the safety-relevant pipeline
- ✓How reinforcement learning is used for post-training and alignment, from a lecture with John Schulman
- ✓How frontier models are evaluated for offensive cyber capabilities, with Nicholas Carlini of Anthropic
- ✓What mechanistic interpretability can currently tell you about a model's internals, with Jack Lindsey
- ✓How labs write and enforce model policies, and the tradeoffs open-weight releases introduce
- ✓How to threat-model AI misuse in biosecurity and reason about recursive self-improvement trajectories
- ✓How to design, run and present a small empirical safety experiment on a real model
Highlights
- •Guest lecturers are the people doing the work at frontier labs: John Schulman, Jakub Pachocki, Nicholas Carlini, Jack Lindsey and Nathan Lambert
- •Lectures go up on YouTube: the Fall 2026 Lecture 3 on modern LLM training and inference is already public
- •The Fall 2025 edition is fully public (syllabus, videos, rubrics, student papers and posters), so you can study both years side by side
- •The head TA's public retrospective on the first run flags its weak spots honestly: limited hands-on technical depth, scattered materials and late project guidance
Who It’s For
Best For
- ✓ML engineers working on post-training, evals or red-teaming who want the safety research framing
- ✓Researchers deciding whether to move into alignment, interpretability or AI security
- ✓Technical leads who need a grounded view of frontier-model risk beyond press coverage
- ✓Self-learners who want to replicate published safety experiments on open models
Prerequisites
- •Mathematical maturity: proofs, probability and information theory
- •Undergraduate machine learning at the level of Harvard CS 181 or MIT 6.036
- •Comfortable programming in Python and training a basic neural network
FAQ
What is CS 2881R: AI Safety (Harvard, Fall 2026)?
CS 2881R is Harvard's graduate course on AI safety, taught by Boaz Barak, with lectures published on YouTube. It is for ML engineers and researchers who already know how LLMs are trained. They will learn how frontier labs think about post-training, alignment, interpretability, model policy and misuse risk, and how to run small safety experiments of their own.
Is CS 2881R: AI Safety (Harvard, Fall 2026) free?
CS 2881R: AI Safety (Harvard, Fall 2026) is free to access.
What level is CS 2881R: AI Safety (Harvard, Fall 2026) for?
CS 2881R: AI Safety (Harvard, Fall 2026) is aimed at a advanced audience. Recommended background: Mathematical maturity: proofs, probability and information theory, Undergraduate machine learning at the level of Harvard CS 181 or MIT 6.036, Comfortable programming in Python and training a basic neural network.
How long does CS 2881R: AI Safety (Harvard, Fall 2026) take?
Expect roughly 13 weekly lectures (~2h45m each), Sep 3 – Dec 3, 2026; recordings posted to YouTube, sometimes with a delay. Most learners work through it at their own pace.
What will I learn from CS 2881R: AI Safety (Harvard, Fall 2026)?
You'll learn: How modern LLMs are pre-trained, post-trained and served, as the safety-relevant pipeline; How reinforcement learning is used for post-training and alignment, from a lecture with John Schulman; How frontier models are evaluated for offensive cyber capabilities, with Nicholas Carlini of Anthropic; What mechanistic interpretability can currently tell you about a model's internals, with Jack Lindsey; How labs write and enforce model policies, and the tradeoffs open-weight releases introduce; How to threat-model AI misuse in biosecurity and reason about recursive self-improvement trajectories; How to design, run and present a small empirical safety experiment on a real model.
Topics
Sources
This page was written from 3 sources, 2 on domains other than boazbk.github.io.