A Visual Guide to Attention Variants in Modern LLMs

by Sebastian Raschka (Ahead of AI)

IntermediateGuideFree~45 min read; longer if you run the companion code

MHA, GQA, MLA, sliding-window, sparse, gated and hybrid linear attention, each drawn and mapped to the models that use it.

Start LearningAdded Oct 10, 2026 · Updated Oct 10, 2026

Overview

Published March 22, 2026 in Sebastian Raschka's Ahead of AI newsletter, this guide works through the attention mechanisms found in current open-weight LLMs in seven numbered sections. It starts with multi-head attention: why attention was invented, the masked attention matrix, self-attention internals and the step from one head to many. Grouped-query attention follows, with its KV-cache memory savings and why it still matters in 2026 in models such as Llama 3, Qwen3 and Gemma 3. Multi-head latent attention is presented as compression rather than sharing, with ablation results and its spread from DeepSeek V3 to Kimi K2, GLM-5 and Mistral Large 3. The sliding-window section uses Gemma 3 as a reference for the ratio of local to global layers and the window size, and covers combining SWA with GQA. DeepSeek Sparse Attention is contrasted with sliding windows and paired with MLA. The gated attention section covers the output gate, zero-centered QK-Norm and partial RoPE in Qwen3-Next and Qwen3.5. The hybrid section covers Gated DeltaNet, Kimi Delta Attention, Ling 2.5's Lightning Attention and Nemotron's Mamba-2 layers. Several sections link to runnable implementations in the ch04 folder of Raschka's LLMs-from-scratch GitHub repository (GQA, MLA, SWA, DeltaNet).

At a Glance

Topic
Models
Level
Intermediate
Format
Guide
Cost
Free
Duration
~45 min read; longer if you run the companion code
Provider
Sebastian Raschka (Ahead of AI)
Hands-on
Yes — code/exercises
Certificate
None

What You’ll Learn

  • ✓How masked multi-head self-attention computes queries, keys, values and per-head outputs
  • ✓How grouped-query attention shares key-value heads to shrink the KV cache
  • ✓How multi-head latent attention caches a compressed latent and why DeepSeek-style models adopted it
  • ✓How sliding-window attention trades context reach for memory, using Gemma 3's local-to-global ratio
  • ✓How DeepSeek Sparse Attention uses a learned indexer and top-k selection instead of a fixed window
  • ✓How gated attention adds an output gate, zero-centered QK-Norm and partial RoPE for training stability
  • ✓How hybrid designs replace most attention layers with Gated DeltaNet, Lightning Attention or Mamba-2 blocks

Highlights

  • •Maps every variant to named production models, from Llama 3 and Gemma 3 to DeepSeek V3.2, Kimi Linear and Nemotron 3
  • •Companion from-scratch PyTorch implementations for GQA, MLA, SWA and DeltaNet in the LLMs-from-scratch repo
  • •31 figures, and it ties into Raschka's LLM Architecture Gallery for whole-model diagrams
  • •Discussed on Hacker News (23 points) the day it was published

Who It’s For

Best For

  • ✓Engineers reading open-weight model cards who want to understand the architecture choices
  • ✓Inference engineers estimating KV-cache memory across GQA, MLA and sliding-window models
  • ✓Practitioners who have built a basic GPT and want to implement modern attention variants

Prerequisites

  • •Basic understanding of the transformer architecture and self-attention
  • •Comfort reading PyTorch code if you want to use the companion implementations

FAQ

What is A Visual Guide to Attention Variants in Modern LLMs?

An illustrated reference by Sebastian Raschka for engineers who read model cards and want to know what the attention block in a 2026 LLM actually does. It walks from multi-head attention through GQA, MLA, sliding-window, DeepSeek Sparse Attention, gated attention and hybrid linear attention, with the models that use each and companion PyTorch code.

Is A Visual Guide to Attention Variants in Modern LLMs free?

A Visual Guide to Attention Variants in Modern LLMs is free to access.

What level is A Visual Guide to Attention Variants in Modern LLMs for?

A Visual Guide to Attention Variants in Modern LLMs is aimed at a intermediate audience. Recommended background: Basic understanding of the transformer architecture and self-attention, Comfort reading PyTorch code if you want to use the companion implementations.

How long does A Visual Guide to Attention Variants in Modern LLMs take?

Expect roughly ~45 min read; longer if you run the companion code. Most learners work through it at their own pace.

What will I learn from A Visual Guide to Attention Variants in Modern LLMs?

You'll learn: How masked multi-head self-attention computes queries, keys, values and per-head outputs; How grouped-query attention shares key-value heads to shrink the KV cache; How multi-head latent attention caches a compressed latent and why DeepSeek-style models adopted it; How sliding-window attention trades context reach for memory, using Gemma 3's local-to-global ratio; How DeepSeek Sparse Attention uses a learned indexer and top-k selection instead of a fixed window; How gated attention adds an output gate, zero-centered QK-Norm and partial RoPE for training stability; How hybrid designs replace most attention layers with Gated DeltaNet, Lightning Attention or Mamba-2 blocks.

Topics

attentionGQAMLAsliding window attentionlinear attentionLLM architecture

Sources

This page was written from 3 sources, 2 on domains other than magazine.sebastianraschka.com.

  1. 1.magazine.sebastianraschka.com — visual attention variantsvendor
  2. 2.github.com — ch04
  3. 3.hn.algolia.com — search