MLModelsAgentic

MIT MAS.S60 / 6.S985: Modeling — MultiModal AI (Spring 2026)

by MIT Media Lab (Paul Liang)

AdvancedCourseFree24 lectures of 90 minutes (Spring 2026); 13 recorded on YouTube, self-paced replay with slides and readings

MIT's graduate course on multimodal AI: fusion, alignment, large multimodal models, generation, reasoning and multimodal agents.

Start LearningAdded Sep 17, 2026 · Updated Sep 17, 2026

Overview

MAS.S60 / 6.S985, Modeling: MultiModal AI, is the Spring 2026 edition of the MIT Media Lab course subtitled "How to AI (Almost) Anything", taught by Paul Liang with Dimitris Bertsimas, Jinhua Zhao and Sang-Gook Kim. It ran from February 3 to May 12, 2026, with two 90-minute lectures a week. The schedule opens with the course introduction, multimodal datasets, an AI tutorial on training neural networks and PyTorch, and data heterogeneity (geometric deep learning, vision transformers, attention). It then covers multimodal fusion and advanced fusion (quantifying cross-modal interactions, Kosmos-2 and Chameleon), alignment (contrastive learning and vision-language models), and large multimodal models (compute-optimal training, LoRA, mixture-of-experts, quantization, visual instruction tuning), with a public Colab tutorial on multimodal LLMs. Next come multimodal generation and modern generative AI (VAEs, diffusion and flow models, controllable generation), a midterm, multimodal reasoning (reinforcement learning from preferences, reasoning LLMs), explainable reasoning, multimodal interaction (GUI agents, web task evaluation, RLHF) and cross-modal transfer. Application sessions cover manufacturing, design, prescriptive modeling, cities and transportation, alongside an agents tutorial on multimodal agent pipelines and evaluation. The course closes with self-evolving AI and "AI for new senses" such as touch, smell and taste. Every lecture has a PDF slide deck and a reading list on the course site, and 13 sessions have YouTube recordings. Enrolled students also completed homework and a team research project. The course's project page lists earlier student work published at EMNLP, NeurIPS, ICCV and ACL. The spring 2025 edition is on MIT OpenCourseWare as a graduate-level course.

At a Glance

Topic
ML
Level
Advanced
Format
Course
Cost
Free
Duration
24 lectures of 90 minutes (Spring 2026); 13 recorded on YouTube, self-paced replay with slides and readings
Provider
MIT Media Lab (Paul Liang)
Hands-on
No
Certificate
None

What You’ll Learn

  • Represent heterogeneous data modalities using attention, vision transformers and geometric deep learning
  • Compare multimodal fusion strategies and quantify the interactions between different input modalities
  • Align vision and language with contrastive learning as used in modern vision-language models
  • Explain how large multimodal models use LoRA, mixture-of-experts, quantization and visual instruction tuning
  • Describe VAEs, diffusion and flow models for controllable multimodal content generation
  • Analyze multimodal reasoning, preference-based reinforcement learning, and GUI and web agents
  • Apply cross-modal transfer, co-learning and self-training when a modality lacks labeled data
  • Design and evaluate multimodal agent pipelines following the course's dedicated agents tutorial

Highlights

  • Every lecture ships a public PDF slide deck and a reading list, so the site doubles as an annotated multimodal-ML bibliography
  • 13 recorded YouTube lectures cover fusion, alignment, large multimodal models, generation, reasoning, interaction and self-evolving AI
  • Goes beyond vision and language to sensors, manufacturing, cities, transportation and even touch, smell and taste
  • A public Colab tutorial on multimodal LLMs offers a hands-on entry point to otherwise lecture-based material
  • Earlier student projects from the course have been published at EMNLP, NeurIPS, ICCV and ACL

Who It’s For

Best For

  • ML engineers moving from text-only LLMs to vision-language and multimodal systems
  • Graduate students and researchers scoping a multimodal research project
  • Engineers fusing sensor, industrial or urban data who want principled fusion and alignment methods
  • Builders of GUI and multimodal agents who want the theory underneath them

Prerequisites

  • Solid deep learning fundamentals; MIT OpenCourseWare lists the course at graduate level
  • Working Python; the course's own tutorial lecture introduces PyTorch and neural-network training
  • Comfort reading research papers, since every lecture is built around a reading list

FAQ

What is MIT MAS.S60 / 6.S985: Modeling — MultiModal AI (Spring 2026)?

A graduate MIT course on multimodal AI: how to represent, fuse, align, reason over and generate across language, vision, audio, sensors and other data types. It is for ML engineers and researchers who already know deep learning and want the principles behind vision-language models, multimodal LLMs and multimodal agents, using public slides, reading lists and 13 recorded lectures.

Is MIT MAS.S60 / 6.S985: Modeling — MultiModal AI (Spring 2026) free?

MIT MAS.S60 / 6.S985: Modeling — MultiModal AI (Spring 2026) is free to access.

What level is MIT MAS.S60 / 6.S985: Modeling — MultiModal AI (Spring 2026) for?

MIT MAS.S60 / 6.S985: Modeling — MultiModal AI (Spring 2026) is aimed at a advanced audience. Recommended background: Solid deep learning fundamentals; MIT OpenCourseWare lists the course at graduate level, Working Python; the course's own tutorial lecture introduces PyTorch and neural-network training, Comfort reading research papers, since every lecture is built around a reading list.

How long does MIT MAS.S60 / 6.S985: Modeling — MultiModal AI (Spring 2026) take?

Expect roughly 24 lectures of 90 minutes (Spring 2026); 13 recorded on YouTube, self-paced replay with slides and readings. Most learners work through it at their own pace.

What will I learn from MIT MAS.S60 / 6.S985: Modeling — MultiModal AI (Spring 2026)?

You'll learn: Represent heterogeneous data modalities using attention, vision transformers and geometric deep learning; Compare multimodal fusion strategies and quantify the interactions between different input modalities; Align vision and language with contrastive learning as used in modern vision-language models; Explain how large multimodal models use LoRA, mixture-of-experts, quantization and visual instruction tuning; Describe VAEs, diffusion and flow models for controllable multimodal content generation; Analyze multimodal reasoning, preference-based reinforcement learning, and GUI and web agents; Apply cross-modal transfer, co-learning and self-training when a modality lacks labeled data; Design and evaluate multimodal agent pipelines following the course's dedicated agents tutorial.

Topics

multimodal AIvision-language modelsmultimodal LLMsdiffusion modelsMIT courserepresentation learning

Sources

This page was written from 5 sources, 2 on domains other than mit-mi.github.io.

  1. 1.mit-mi.github.iospring2026vendor
  2. 2.mit-mi.github.ioschedulevendor
  3. 3.mit-mi.github.ioprojectsvendor
  4. 4.heyuan110.com2026 09 09 free ai agent courses fall 2026
  5. 5.ocw.mit.edumas s60 how to ai almost anything spring 2025