CS25: Transformers United V6
by Stanford University
Nine frontier researchers explain the architecture work they shipped this year — free, no signup, no Stanford affiliation.
Overview
CS25 is a Stanford seminar course, not a tutorial series: each week a working researcher presents current work, and the V6 edition ran Spring quarter 2026 from March 30 to June 3, hosted by Steven Feng, Karan P. Singh, Michael C. Frank and Christopher Manning. The nine-lecture schedule opens with the instructors' own overview of transformer history and trends, then hands the room to Hazel Nam and Lucas Maes on joint embedding predictive architectures and latent-space world modeling, Albert Gu of CMU and Cartesia AI on the tradeoffs between state space models and transformers, Nouamane Tazi of Hugging Face on scaling training to thousands of GPUs, Shrimai Prabhumoye of Mistral AI on the future of pretraining, Andrew Lampinen of Anthropic on how generalization differs when it comes from parameters versus from context, Vivek Natarajan of DeepMind on collaborative agents for science and medicine, Victoria Lin of Thinking Machines on natively multimodal intelligence, and Charles Frye of Modal on serving transformers in production. There is no coursework beyond attendance, no problem sets and no certificate; the value is that the people who built Mamba, the Ultra-Scale Playbook and Modal's inference stack explain their own results and take questions. Anyone can audit in person, join the Zoom livestream, or watch the recordings afterwards without signing up or being affiliated with Stanford, and V1 through V5 remain archived on the same site.
At a Glance
- Topic
- Models
- Level
- Advanced
- Format
- Video
- Cost
- Free
- Duration
- 9 lectures, ~75 min each (~11 hours total), self-paced via recordings
- Provider
- Stanford University
- Hands-on
- No
- Certificate
- None
What You’ll Learn
- ✓Trace how attention displaced recurrence and where transformer research is heading next
- ✓Explain joint embedding predictive architectures and the shift to latent-space world modeling
- ✓Compare state space models against transformers on the tradeoffs Albert Gu identifies
- ✓Scale a training run to thousands of GPUs using Hugging Face's ultra-scale methods
- ✓Understand where pretraining is going beyond straightforward next-token prediction objectives
- ✓Distinguish generalization that comes from model parameters versus from in-context evidence
- ✓Design collaborative multi-agent systems aimed at scientific and medical research problems
- ✓Apply production inference lessons from Modal on serving transformers under real load
Highlights
- •Speakers are the practitioners behind the work: Albert Gu on state space models, Nouamane Tazi on ultra-scale training, Charles Frye on production inference, Vivek Natarajan on research agents
- •Sixth iteration of the seminar, with V1 through V5 archived on the same site — five years of frontier talks in one place
- •Genuinely open: no signup, no Stanford affiliation, no fee; audit in person, join the Zoom livestream, or watch the recordings
- •A research seminar rather than a tutorial — each session is current work presented by its author, with live questions, not a polished course module
- •Active Discord community of 5,000+ members discussing the lectures
Who It’s For
Best For
- ✓AI engineers who ship with LLMs and want the architecture reasoning underneath the API
- ✓ML practitioners tracking state space models, world models and multimodal pretraining
- ✓Infrastructure and inference engineers sizing up serving and large-scale training decisions
- ✓Researchers and grad students who want a survey of 2026 frontier work from the authors
Prerequisites
- •Solid grasp of the transformer architecture — attention, positional encoding, pretraining vs. fine-tuning
- •Comfort reading ML research papers; talks assume familiarity with current literature and terminology
- •Undergraduate linear algebra and probability; no coding assignments, so no implementation experience required
FAQ
What is CS25: Transformers United V6?
Stanford's flagship transformers seminar, now in its sixth iteration, brings a different frontier researcher to the podium each week to present work they published this year. It is aimed at engineers who already build with LLMs and want the architectural reasoning underneath them: state space model tradeoffs, latent-space world models, multimodal pretraining, and what production inference actually costs. Free, open to anyone, and recorded.
Is CS25: Transformers United V6 free?
CS25: Transformers United V6 is free to access.
What level is CS25: Transformers United V6 for?
CS25: Transformers United V6 is aimed at a advanced audience. Recommended background: Solid grasp of the transformer architecture — attention, positional encoding, pretraining vs. fine-tuning, Comfort reading ML research papers; talks assume familiarity with current literature and terminology, Undergraduate linear algebra and probability; no coding assignments, so no implementation experience required.
How long does CS25: Transformers United V6 take?
Expect roughly 9 lectures, ~75 min each (~11 hours total), self-paced via recordings. Most learners work through it at their own pace.
What will I learn from CS25: Transformers United V6?
You'll learn: Trace how attention displaced recurrence and where transformer research is heading next; Explain joint embedding predictive architectures and the shift to latent-space world modeling; Compare state space models against transformers on the tradeoffs Albert Gu identifies; Scale a training run to thousands of GPUs using Hugging Face's ultra-scale methods; Understand where pretraining is going beyond straightforward next-token prediction objectives; Distinguish generalization that comes from model parameters versus from in-context evidence; Design collaborative multi-agent systems aimed at scientific and medical research problems; Apply production inference lessons from Modal on serving transformers under real load.
Topics
Sources
This page was written from 3 sources, 2 on domains other than web.stanford.edu.