AgenticModelsML

CS2680: Modern AI Systems: Agents and System Optimizations (Harvard, Fall 2026)

by Harvard SEAS (Juncheng Yang)

AdvancedCourseFree14-week semester (September 2 to early December 2026), two 75-minute lectures a week; slides are posted as the term runs

A Harvard graduate course that follows one agent from Claude Code usage and agent design down to KV-cache, quantization and speculative decoding in the serving stack

Start LearningAdded Oct 3, 2026 · Updated Oct 3, 2026

Overview

CS2680, Modern AI Systems: Agents and System Optimizations, is taught at Harvard SEAS in Fall 2026 by Juncheng Yang, an assistant professor whose cache-eviction algorithms SIEVE (NSDI'24) and S3-FIFO (SOSP'23) are deployed in production systems, and who runs FreeInference, a free LLM inference service. The course has two parts. Part I, Introduction to LLMs and agents, opens with how Claude Code works and multi-agent workflows, then covers agent design: context management and tool design, multi-agent architecture and communication, agent memory and reliability, agent evaluation and cost-effective agents, and agent safety with self-evolving harnesses, before a lecture on large language models. Part II, Systems for LLMs and agents, covers efficient serving: paging, batching and scheduling, KV-cache optimization, pruning and quantization, speculative decoding, and routing and load balancing for agentic systems. Five assignments build one stack: Mads-Lens (observe an agent), Mads-Loop (build the loop), Mads-Opt (token optimization), Mads-Serve (serve a model) and Mads-Stack (full-stack optimization), followed by a final project with a poster and demo. Public reference notes include an agent-evaluation guide that runs from GLUE and MMLU to SWE-bench, Terminal-Bench, OSWorld, tau-bench and BFCL. Slides for lectures 1 to 6 were public at the start of October 2026. The API access and the 120-GPU RTX PRO 6000 Blackwell cluster are for enrolled students only, so self-learners bring their own compute.

At a Glance

Topic
Agentic
Level
Advanced
Format
Course
Cost
Free
Duration
14-week semester (September 2 to early December 2026), two 75-minute lectures a week; slides are posted as the term runs
Provider
Harvard SEAS (Juncheng Yang)
Hands-on
Yes — code/exercises
Certificate
None

What You’ll Learn

  • ✓Explain how Claude Code runs its agent loop and use multi-agent workflows for software engineering
  • ✓Design agent context management and tools, and choose between multi-agent architectures and communication patterns
  • ✓Add memory and reliability mechanisms to an agent and evaluate it for both quality and cost
  • ✓Pick agent benchmarks such as SWE-bench, Terminal-Bench, OSWorld, tau-bench and BFCL for a given capability
  • ✓Serve an open-weight model and tune paging, batching and request scheduling for throughput
  • ✓Apply KV-cache and prefix reuse, quantization and speculative decoding to cut serving latency and cost
  • ✓Route and load-balance agent traffic across model backends as a full-stack optimization problem

Highlights

  • •Joins agent design and LLM serving systems in one course, so token-level agent costs connect directly to KV-cache and scheduling choices
  • •The five Mads assignments build on each other, from watching an agent work to optimizing the full stack under it
  • •Taught by a systems researcher whose caching algorithms run in production at Google, VMware and Redpanda
  • •The public agent-evaluation reference note is a curated map of benchmarks from GLUE through Terminal-Bench and ARC-AGI-3, with links to papers and leaderboards

Who It’s For

Best For

  • ✓Engineers running agents in production who need to cut latency and inference cost
  • ✓Systems and infrastructure engineers moving into LLM serving and agent platforms
  • ✓Graduate students looking for a current agents-plus-systems syllabus to self-study alongside the term

Prerequisites

  • •An undergraduate computer systems course (Harvard's CS61) and at least one graduate-level systems course
  • •Comfort with Python and PyTorch
  • •Your own model API access and a GPU if you want to do the assignments outside Harvard

FAQ

What is CS2680: Modern AI Systems: Agents and System Optimizations (Harvard, Fall 2026)?

CS2680 is a Fall 2026 Harvard graduate course for systems-minded engineers who build LLM agents and want to understand what happens under them. Its public site posts the schedule, slides and reference notes. You learn to design, measure and optimize an agent, then serve an open-weight model yourself and tune batching, caching, quantization and decoding.

Is CS2680: Modern AI Systems: Agents and System Optimizations (Harvard, Fall 2026) free?

CS2680: Modern AI Systems: Agents and System Optimizations (Harvard, Fall 2026) is free to access.

What level is CS2680: Modern AI Systems: Agents and System Optimizations (Harvard, Fall 2026) for?

CS2680: Modern AI Systems: Agents and System Optimizations (Harvard, Fall 2026) is aimed at a advanced audience. Recommended background: An undergraduate computer systems course (Harvard's CS61) and at least one graduate-level systems course, Comfort with Python and PyTorch, Your own model API access and a GPU if you want to do the assignments outside Harvard.

How long does CS2680: Modern AI Systems: Agents and System Optimizations (Harvard, Fall 2026) take?

Expect roughly 14-week semester (September 2 to early December 2026), two 75-minute lectures a week; slides are posted as the term runs. Most learners work through it at their own pace.

What will I learn from CS2680: Modern AI Systems: Agents and System Optimizations (Harvard, Fall 2026)?

You'll learn: Explain how Claude Code runs its agent loop and use multi-agent workflows for software engineering; Design agent context management and tools, and choose between multi-agent architectures and communication patterns; Add memory and reliability mechanisms to an agent and evaluate it for both quality and cost; Pick agent benchmarks such as SWE-bench, Terminal-Bench, OSWorld, tau-bench and BFCL for a given capability; Serve an open-weight model and tune paging, batching and request scheduling for throughput; Apply KV-cache and prefix reuse, quantization and speculative decoding to cut serving latency and cost; Route and load-balance agent traffic across model backends as a full-stack optimization problem.

Topics

AI agentsLLM servingKV cachespeculative decodingagent evaluationClaude Code

Sources

This page was written from 5 sources, 3 on domains other than cs2680.com.

  1. 1.cs2680.com — cs2680.comvendor
  2. 2.course.agentic-system.org — course.agentic-system.org
  3. 3.cs2680.com — agent evaluationvendor
  4. 4.junchengyang.com — teaching
  5. 5.juncheng.seas.harvard.edu — juncheng.seas.harvard.edu