ModelsMLFine-TuningAgentic

CS 2881R: AI Safety (Harvard, Fall 2026)

by Harvard University (Boaz Barak)

AdvancedCourseFree13 weekly lectures (~2h45m each), Sep 3 – Dec 3, 2026; recordings posted to YouTube, sometimes with a delay

Harvard's graduate AI safety seminar, 2026 edition: RL post-training, interpretability, cyber and bio risk, taught with guests from Anthropic and OpenAI.

Start LearningAdded Sep 25, 2026 · Updated Sep 25, 2026

Overview

CS 2881R: AI Safety is a Harvard computer science course taught by Boaz Barak, a Harvard CS professor. The Fall 2026 offering is its second run, after the first in Fall 2025. It meets weekly from September 3 to December 3, 2026, for 13 in-person lectures of about 2 hours 45 minutes each. In each session a group of students presents an experiment, and the class discusses it. The public schedule runs from an introduction to cyber capabilities (guest Nicholas Carlini of Anthropic) and then modern LLM training and inference. Next come recursive self-improvement and AI trajectories (Dwarkesh Patel, Daniel Kokotajlo) and the economic impact of AI (Chad Jones, Erik Brynjolfsson). Later sessions cover reinforcement learning for post-training and alignment (John Schulman), model policies (Ziad Reslan), open-source models (Nathan Lambert), AI interpretability (Jack Lindsey), alignment in the age of recursive self-improvement (Jakub Pachocki), and AI biosecurity and threat modeling (Luca Righetti). Two slots are still TBD. Enrolled students give experiment presentations, write scribe notes and do a final project. The 2025 edition added a paper-replication midterm and a NeurIPS-style final paper. The course encourages heavy use of generative and agentic AI for coursework. The syllabus says lectures will be recorded and published online where technically possible, and Lecture 3 (Modern LLM Training and Inference) is already on YouTube. The 2025 syllabus, videos, rubrics and student papers are public too.

At a Glance

Topic
Models
Level
Advanced
Format
Course
Cost
Free
Duration
13 weekly lectures (~2h45m each), Sep 3 – Dec 3, 2026; recordings posted to YouTube, sometimes with a delay
Provider
Harvard University (Boaz Barak)
Hands-on
Yes — code/exercises
Certificate
None

What You’ll Learn

  • ✓How modern LLMs are pre-trained, post-trained and served, as the safety-relevant pipeline
  • ✓How reinforcement learning is used for post-training and alignment, from a lecture with John Schulman
  • ✓How frontier models are evaluated for offensive cyber capabilities, with Nicholas Carlini of Anthropic
  • ✓What mechanistic interpretability can currently tell you about a model's internals, with Jack Lindsey
  • ✓How labs write and enforce model policies, and the tradeoffs open-weight releases introduce
  • ✓How to threat-model AI misuse in biosecurity and reason about recursive self-improvement trajectories
  • ✓How to design, run and present a small empirical safety experiment on a real model

Highlights

  • •Guest lecturers are the people doing the work at frontier labs: John Schulman, Jakub Pachocki, Nicholas Carlini, Jack Lindsey and Nathan Lambert
  • •Lectures go up on YouTube: the Fall 2026 Lecture 3 on modern LLM training and inference is already public
  • •The Fall 2025 edition is fully public (syllabus, videos, rubrics, student papers and posters), so you can study both years side by side
  • •The head TA's public retrospective on the first run flags its weak spots honestly: limited hands-on technical depth, scattered materials and late project guidance

Who It’s For

Best For

  • ✓ML engineers working on post-training, evals or red-teaming who want the safety research framing
  • ✓Researchers deciding whether to move into alignment, interpretability or AI security
  • ✓Technical leads who need a grounded view of frontier-model risk beyond press coverage
  • ✓Self-learners who want to replicate published safety experiments on open models

Prerequisites

  • •Mathematical maturity: proofs, probability and information theory
  • •Undergraduate machine learning at the level of Harvard CS 181 or MIT 6.036
  • •Comfortable programming in Python and training a basic neural network

FAQ

What is CS 2881R: AI Safety (Harvard, Fall 2026)?

CS 2881R is Harvard's graduate course on AI safety, taught by Boaz Barak, with lectures published on YouTube. It is for ML engineers and researchers who already know how LLMs are trained. They will learn how frontier labs think about post-training, alignment, interpretability, model policy and misuse risk, and how to run small safety experiments of their own.

Is CS 2881R: AI Safety (Harvard, Fall 2026) free?

CS 2881R: AI Safety (Harvard, Fall 2026) is free to access.

What level is CS 2881R: AI Safety (Harvard, Fall 2026) for?

CS 2881R: AI Safety (Harvard, Fall 2026) is aimed at a advanced audience. Recommended background: Mathematical maturity: proofs, probability and information theory, Undergraduate machine learning at the level of Harvard CS 181 or MIT 6.036, Comfortable programming in Python and training a basic neural network.

How long does CS 2881R: AI Safety (Harvard, Fall 2026) take?

Expect roughly 13 weekly lectures (~2h45m each), Sep 3 – Dec 3, 2026; recordings posted to YouTube, sometimes with a delay. Most learners work through it at their own pace.

What will I learn from CS 2881R: AI Safety (Harvard, Fall 2026)?

You'll learn: How modern LLMs are pre-trained, post-trained and served, as the safety-relevant pipeline; How reinforcement learning is used for post-training and alignment, from a lecture with John Schulman; How frontier models are evaluated for offensive cyber capabilities, with Nicholas Carlini of Anthropic; What mechanistic interpretability can currently tell you about a model's internals, with Jack Lindsey; How labs write and enforce model policies, and the tradeoffs open-weight releases introduce; How to threat-model AI misuse in biosecurity and reason about recursive self-improvement trajectories; How to design, run and present a small empirical safety experiment on a real model.

Topics

ai safetyalignmentinterpretabilityrl post-trainingharvardred teaming

Sources

This page was written from 3 sources, 2 on domains other than boazbk.github.io.

  1. 1.boazbk.github.io — mltheoryseminarvendor
  2. 2.youtube.com — watch
  3. 3.lesswrong.com — reflections on ta ing harvard s first ai safety course