A Visual Guide to Attention Variants in Modern LLMs
by Sebastian Raschka (Ahead of AI)
MHA, GQA, MLA, sliding-window, sparse, gated and hybrid linear attention, each drawn and mapped to the models that use it.
Overview
Published March 22, 2026 in Sebastian Raschka's Ahead of AI newsletter, this guide works through the attention mechanisms found in current open-weight LLMs in seven numbered sections. It starts with multi-head attention: why attention was invented, the masked attention matrix, self-attention internals and the step from one head to many. Grouped-query attention follows, with its KV-cache memory savings and why it still matters in 2026 in models such as Llama 3, Qwen3 and Gemma 3. Multi-head latent attention is presented as compression rather than sharing, with ablation results and its spread from DeepSeek V3 to Kimi K2, GLM-5 and Mistral Large 3. The sliding-window section uses Gemma 3 as a reference for the ratio of local to global layers and the window size, and covers combining SWA with GQA. DeepSeek Sparse Attention is contrasted with sliding windows and paired with MLA. The gated attention section covers the output gate, zero-centered QK-Norm and partial RoPE in Qwen3-Next and Qwen3.5. The hybrid section covers Gated DeltaNet, Kimi Delta Attention, Ling 2.5's Lightning Attention and Nemotron's Mamba-2 layers. Several sections link to runnable implementations in the ch04 folder of Raschka's LLMs-from-scratch GitHub repository (GQA, MLA, SWA, DeltaNet).
At a Glance
- Topic
- Models
- Level
- Intermediate
- Format
- Guide
- Cost
- Free
- Duration
- ~45 min read; longer if you run the companion code
- Provider
- Sebastian Raschka (Ahead of AI)
- Hands-on
- Yes — code/exercises
- Certificate
- None
What You’ll Learn
- ✓How masked multi-head self-attention computes queries, keys, values and per-head outputs
- ✓How grouped-query attention shares key-value heads to shrink the KV cache
- ✓How multi-head latent attention caches a compressed latent and why DeepSeek-style models adopted it
- ✓How sliding-window attention trades context reach for memory, using Gemma 3's local-to-global ratio
- ✓How DeepSeek Sparse Attention uses a learned indexer and top-k selection instead of a fixed window
- ✓How gated attention adds an output gate, zero-centered QK-Norm and partial RoPE for training stability
- ✓How hybrid designs replace most attention layers with Gated DeltaNet, Lightning Attention or Mamba-2 blocks
Highlights
- •Maps every variant to named production models, from Llama 3 and Gemma 3 to DeepSeek V3.2, Kimi Linear and Nemotron 3
- •Companion from-scratch PyTorch implementations for GQA, MLA, SWA and DeltaNet in the LLMs-from-scratch repo
- •31 figures, and it ties into Raschka's LLM Architecture Gallery for whole-model diagrams
- •Discussed on Hacker News (23 points) the day it was published
Who It’s For
Best For
- ✓Engineers reading open-weight model cards who want to understand the architecture choices
- ✓Inference engineers estimating KV-cache memory across GQA, MLA and sliding-window models
- ✓Practitioners who have built a basic GPT and want to implement modern attention variants
Prerequisites
- •Basic understanding of the transformer architecture and self-attention
- •Comfort reading PyTorch code if you want to use the companion implementations
FAQ
What is A Visual Guide to Attention Variants in Modern LLMs?
An illustrated reference by Sebastian Raschka for engineers who read model cards and want to know what the attention block in a 2026 LLM actually does. It walks from multi-head attention through GQA, MLA, sliding-window, DeepSeek Sparse Attention, gated attention and hybrid linear attention, with the models that use each and companion PyTorch code.
Is A Visual Guide to Attention Variants in Modern LLMs free?
A Visual Guide to Attention Variants in Modern LLMs is free to access.
What level is A Visual Guide to Attention Variants in Modern LLMs for?
A Visual Guide to Attention Variants in Modern LLMs is aimed at a intermediate audience. Recommended background: Basic understanding of the transformer architecture and self-attention, Comfort reading PyTorch code if you want to use the companion implementations.
How long does A Visual Guide to Attention Variants in Modern LLMs take?
Expect roughly ~45 min read; longer if you run the companion code. Most learners work through it at their own pace.
What will I learn from A Visual Guide to Attention Variants in Modern LLMs?
You'll learn: How masked multi-head self-attention computes queries, keys, values and per-head outputs; How grouped-query attention shares key-value heads to shrink the KV cache; How multi-head latent attention caches a compressed latent and why DeepSeek-style models adopted it; How sliding-window attention trades context reach for memory, using Gemma 3's local-to-global ratio; How DeepSeek Sparse Attention uses a learned indexer and top-k selection instead of a fixed window; How gated attention adds an output gate, zero-centered QK-Norm and partial RoPE for training stability; How hybrid designs replace most attention layers with Gated DeltaNet, Lightning Attention or Mamba-2 blocks.
Topics
Sources
This page was written from 3 sources, 2 on domains other than magazine.sebastianraschka.com.