Topic

vendor benchmarks

Every THE D[AI]LY BRIEF article on vendor benchmarks — enterprise AI analysis, benchmarks, vendor comparisons, and ROI frameworks for technology and business leaders. Updated as new coverage publishes.

KV cache

Random Eviction Matched the Scorers. Protect the Prompt.

Salesforce AI Research and UIUC replaced the scoring pass in KV cache eviction with a uniform random draw, kept the prompt pinned, and matched the strongest published evictor across four models and six reasoning tasks — at 32-43% higher vLLM throughput. Random-plus-prompt-protection is now the baseline any 'intelligent cache compression' claim has to beat.

September 9, 2026 · 13 min read
AI inference

3,400 Tokens/s Was Batch 1. At 100K Context, It's Batch 12.

Nvidia's 3,400 tokens/sec on Groq 3 LPX and Cerebras' 4,400 on CS-4 are both single-stream figures. On a 256-LPU rack at 100K context, the memory math caps concurrency near a batch of 12 — so any capacity plan sized off a headline token rate is sized for one user.

August 29, 2026 · 12 min read
Qwen3.8-Max

Alibaba's Agent Coded 16 Days. A Human Wrote 13 Commits.

Alibaba says Qwen3.8-Max coded unattended for 16 days, and unlike every rival long-horizon claim it shipped the whole commit trace on public GitHub. Open it and the 16 days become 'more than ten' in Alibaba's own words, 13 of 648 commits turn out to be a human's, and one person still holds the merge keys.

August 5, 2026 · 10 min read