10-423/10-623/10-723: Generative AI (CMU, Fall 2026)
by Carnegie Mellon University (Matt Gormley & Aran Nayebi)
CMU's full-semester generative AI course, from RNN language models to diffusion, MoE, reasoning models and agents
Overview
10-423/10-623/10-723 Generative AI is Carnegie Mellon's semester-long course on the techniques behind modern generative models and foundation models. The Fall 2026 offering is taught by Matt Gormley, who has run every edition since Spring 2024, and Aran Nayebi. The public course site (mlcourse.org/10423) lists 26 lectures that start with RNN language models and automatic differentiation and move through transformer language models, decoding, pre-training and fine-tuning, and modern transformer variants. The course then turns to vision (CNNs, Vision Transformers), variational autoencoders, diffusion models and score matching, GANs, normalizing flows and flow matching. The second half covers parameter-efficient fine-tuning, in-context learning, instruction tuning and RLHF, Direct Preference Optimization, latent diffusion, vision-language models, diffusion transformers, scaling laws, mixture of experts, distributed training, FlashAttention and efficient decoding, long context and RAG, reasoning models, state space and hybrid models, safety, audio, code generation and autonomous agents, video generation and interactive world models. Assignments run from a PyTorch primer (HW0) through large language models, image generation, applying and adapting LLMs, and multimodal foundation models. According to the syllabus, homeworks mix written questions with programming that asks students to implement core algorithms from scratch and to apply existing libraries to real problems. Graded work also includes six quizzes, two programming tests, one exam and a project. Prerequisites are strict: a working knowledge of introductory ML or deep learning (for example CMU 10-601 or 11-785). Readings are free online. Lectures are recorded to Panopto, but the syllabus does not say whether non-CMU learners can watch them. Slides from earlier offerings are publicly hosted on Gormley's CMU page.
At a Glance
- Topic
- ML
- Level
- Intermediate
- Format
- Course
- Cost
- Free
- Duration
- One semester (Fall 2026): 26 lectures of 80 min plus recitations, 5-6 homeworks and a project
- Provider
- Carnegie Mellon University (Matt Gormley & Aran Nayebi)
- Hands-on
- Yes — code/exercises
- Certificate
- None
What You’ll Learn
- ✓Implement transformer language models and train them with automatic differentiation in PyTorch
- ✓Build and train diffusion models, including score matching, latent diffusion and diffusion transformers
- ✓Compare VAEs, GANs, normalizing flows and flow matching as alternative generative modeling approaches
- ✓Adapt pretrained models with parameter-efficient fine-tuning, instruction tuning, RLHF and Direct Preference Optimization
- ✓Scale training using scaling laws, mixture of experts, distributed training and FlashAttention
- ✓Apply long-context techniques, retrieval-augmented generation and reasoning models to real problems
- ✓Understand vision-language models, audio synthesis, video generation and interactive world models
- ✓Identify real-world failure modes of generative models, including bias, hallucination and safety risks
Highlights
- •Covers text, image, audio and video generation in one course; most LLM courses stop at text
- •Fall 2026 schedule includes current topics: reasoning models, state space and hybrid models, MoE, agents and world models
- •Homeworks cover both building core algorithms from scratch and applying production libraries
- •Taught by Matt Gormley, who has run the course in five consecutive semesters since Spring 2024, so the material is refined rather than new
- •Prior-semester slide decks are publicly hosted on cs.cmu.edu, so outside learners can follow the lecture sequence
Who It’s For
Best For
- ✓ML engineers who want to understand the foundation models they deploy
- ✓Developers moving from applying LLMs to training and adapting generative models
- ✓Self-learners after a structured university curriculum across LLMs and diffusion
- ✓Graduate students preparing for research in generative modeling
Prerequisites
- •Working knowledge of introductory machine learning or deep learning (e.g. CMU 10-601, 10-701 or 11-785)
- •Comfortable Python programming; the course starts with a PyTorch primer
- •Linear algebra, probability and calculus at undergraduate level
FAQ
What is 10-423/10-623/10-723: Generative AI (CMU, Fall 2026)?
Carnegie Mellon's Fall 2026 Generative AI course covers the machine learning behind modern foundation models. It is for engineers who already know introductory ML or deep learning and want to implement transformers and diffusion models themselves, then adapt, scale and apply them to text, code, images, audio and video.
Is 10-423/10-623/10-723: Generative AI (CMU, Fall 2026) free?
10-423/10-623/10-723: Generative AI (CMU, Fall 2026) is free to access.
What level is 10-423/10-623/10-723: Generative AI (CMU, Fall 2026) for?
10-423/10-623/10-723: Generative AI (CMU, Fall 2026) is aimed at a intermediate audience. Recommended background: Working knowledge of introductory machine learning or deep learning (e.g. CMU 10-601, 10-701 or 11-785), Comfortable Python programming; the course starts with a PyTorch primer, Linear algebra, probability and calculus at undergraduate level.
How long does 10-423/10-623/10-723: Generative AI (CMU, Fall 2026) take?
Expect roughly One semester (Fall 2026): 26 lectures of 80 min plus recitations, 5-6 homeworks and a project. Most learners work through it at their own pace.
What will I learn from 10-423/10-623/10-723: Generative AI (CMU, Fall 2026)?
You'll learn: Implement transformer language models and train them with automatic differentiation in PyTorch; Build and train diffusion models, including score matching, latent diffusion and diffusion transformers; Compare VAEs, GANs, normalizing flows and flow matching as alternative generative modeling approaches; Adapt pretrained models with parameter-efficient fine-tuning, instruction tuning, RLHF and Direct Preference Optimization; Scale training using scaling laws, mixture of experts, distributed training and FlashAttention; Apply long-context techniques, retrieval-augmented generation and reasoning models to real problems; Understand vision-language models, audio synthesis, video generation and interactive world models; Identify real-world failure modes of generative models, including bias, hallucination and safety risks.
Topics
Sources
This page was written from 4 sources, 1 on domains other than mlcourse.org.
- 1.mlcourse.org — 10423vendor
- 2.mlcourse.org — syllabusvendor
- 3.mlcourse.org — previous offeringsvendor
- 4.cs.cmu.edu — 10423 f25