
3,400 Tokens/s Was Batch 1. At 100K Context, It's Batch 12.
Nvidia's 3,400 tokens/sec on Groq 3 LPX and Cerebras' 4,400 on CS-4 are both single-stream figures. On a 256-LPU rack at 100K context, the memory math caps concurrency near a batch of 12 — so any capacity plan sized off a headline token rate is sized for one user.
August 29, 2026 · 12 min read