Modern GPU Programming for MLSys
by MLC (Tianqi Chen and the MLC community)
An open course that takes you from GPU execution basics to a state-of-the-art Blackwell GEMM and a full Flash Attention 4 kernel in TIRx
Overview
Modern GPU Programming for MLSys was published by the MLC project in July 2026, authored by Tianqi Chen, who co-teaches the 15-442/15-642 Machine Learning Systems course with Zhihao Jia, and the MLC community. The source repository (mlc-ai/modern-gpu-programming-for-mlsys, about 1.3k stars) rebuilds the site through GitHub Actions on every push to main, and a Chinese translation sits under /zh/. The material is in four parts. Part I, Understanding the GPU, covers the GPU execution model, what makes a kernel fast, data layout and its notation, how tensor core data layouts evolved, asynchronous data movement with TMA, the Blackwell tensor core instruction tcgen05.mma, Tensor Memory (TMEM), mbarrier coordination and cluster launch control. Part II introduces TIRx (Tensor IR next), a Python DSL in Apache TVM that keeps hardware-level choices explicit, along with its layout API. Part III takes GEMM from a sequential single-tile kernel through K-loop accumulation and spatial tiling, then TMA async loads, software pipelining and persistent kernels, and finally warp specialization, two-CTA clusters and multi-consumer variants. Part IV builds a complete Flash Attention 4 kernel, including its TMEM layout, barrier protocol, online rescaling, causal masking and grouped-query attention. Appendices cover the TIRx language reference, GPU benchmarking, compiler internals and debugging. The kernels target sm_100a, so you need a B200-class GPU to run them, plus TVM 0.26.0 with tvm.tirx and a CUDA build of PyTorch.
At a Glance
- Topic
- ML
- Level
- Advanced
- Format
- Course
- Cost
- Free
- Duration
- Self-paced; four parts plus appendices, with no official time estimate (running every kernel needs a Blackwell GPU)
- Provider
- MLC (Tianqi Chen and the MLC community)
- Hands-on
- Yes — code/exercises
- Certificate
- None
What You’ll Learn
- ✓Explain the GPU execution model and diagnose what limits a kernel's throughput on modern hardware
- ✓Read and reason about tensor core data layouts and the swizzled layouts Blackwell expects
- ✓Move data asynchronously with TMA and coordinate producer and consumer warps using mbarrier
- ✓Use Blackwell tcgen05.mma tensor core instructions and Tensor Memory (TMEM) inside your own kernels
- ✓Write GPU kernels in TIRx, the Python DSL in Apache TVM, using its layout API
- ✓Take a GEMM from a single-tile version to a pipelined, persistent, warp-specialized kernel with two-CTA clusters
- ✓Implement Flash Attention 4 end to end, with online rescaling, causal masking and grouped-query attention
- ✓Benchmark and debug custom kernels using the course's GPU benchmarking and debugging appendices
Highlights
- •Covers Blackwell-specific features (tcgen05, TMEM, cluster launch control) that most CUDA tutorials, written for Ampere or Hopper, do not reach
- •Ends on a full Flash Attention 4 kernel, the attention kernel behind current high-throughput LLM serving
- •Written by Tianqi Chen, co-instructor of the 15-442/15-642 Machine Learning Systems course, and built on TIRx in Apache TVM, so the code is a real compiler stack you can keep using
- •Open source with a CI-built site, runnable examples and a Chinese edition, so errata are fixed in public
Who It’s For
Best For
- ✓Inference and ML systems engineers who write or tune kernels for LLM serving
- ✓Engineers with Blackwell (B200) access who want to understand what FlashAttention and cuBLAS do on that hardware
- ✓Graduate students in ML systems moving from CUDA basics to production-grade kernels
Prerequisites
- •Comfort with C/CUDA-style GPU programming concepts (threads, warps, shared memory) and Python
- •Working knowledge of matrix multiplication and the attention mechanism in transformers
- •Access to an NVIDIA Blackwell GPU (sm_100a, for example a B200) to run the examples, plus Apache TVM 0.26.0 with tvm.tirx
FAQ
What is Modern GPU Programming for MLSys?
Modern GPU Programming for MLSys is a free online course from the MLC project for ML systems and inference engineers who want to write the kernels that LLM serving depends on. It explains Blackwell-class GPU hardware, then has you build tiled, pipelined and warp-specialized GEMM kernels and a complete Flash Attention 4 kernel.
Is Modern GPU Programming for MLSys free?
Modern GPU Programming for MLSys is free to access.
What level is Modern GPU Programming for MLSys for?
Modern GPU Programming for MLSys is aimed at a advanced audience. Recommended background: Comfort with C/CUDA-style GPU programming concepts (threads, warps, shared memory) and Python, Working knowledge of matrix multiplication and the attention mechanism in transformers, Access to an NVIDIA Blackwell GPU (sm_100a, for example a B200) to run the examples, plus Apache TVM 0.26.0 with tvm.tirx.
How long does Modern GPU Programming for MLSys take?
Expect roughly Self-paced; four parts plus appendices, with no official time estimate (running every kernel needs a Blackwell GPU). Most learners work through it at their own pace.
What will I learn from Modern GPU Programming for MLSys?
You'll learn: Explain the GPU execution model and diagnose what limits a kernel's throughput on modern hardware; Read and reason about tensor core data layouts and the swizzled layouts Blackwell expects; Move data asynchronously with TMA and coordinate producer and consumer warps using mbarrier; Use Blackwell tcgen05.mma tensor core instructions and Tensor Memory (TMEM) inside your own kernels; Write GPU kernels in TIRx, the Python DSL in Apache TVM, using its layout API; Take a GEMM from a single-tile version to a pipelined, persistent, warp-specialized kernel with two-CTA clusters; Implement Flash Attention 4 end to end, with online rescaling, causal masking and grouped-query attention; Benchmark and debug custom kernels using the course's GPU benchmarking and debugging appendices.
Topics
Sources
This page was written from 4 sources, 2 on domains other than mlc.ai.