
Nvidia Skipped the Prefill. The Math Went With It.
Nvidia's cross-model KV cache mapper moves a 32,768-token cache in 277.6ms instead of 6,975.3ms of re-prefill. Broken out by benchmark, GSM8K survived on one of six model pairs — Llama 3.1 8B to 70B fell from 81.12 to 14.78.
August 21, 2026 · 12 min read