Topic

self-hosted inference

Every THE D[AI]LY BRIEF article on self-hosted inference — enterprise AI analysis, benchmarks, vendor comparisons, and ROI frameworks for technology and business leaders. Updated as new coverage publishes.

KV cache

Random Eviction Matched the Scorers. Protect the Prompt.

Salesforce AI Research and UIUC replaced the scoring pass in KV cache eviction with a uniform random draw, kept the prompt pinned, and matched the strongest published evictor across four models and six reasoning tasks — at 32-43% higher vLLM throughput. Random-plus-prompt-protection is now the baseline any 'intelligent cache compression' claim has to beat.

September 9, 2026 · 13 min read
server DRAM

Nvidia Is Losing 3 Points to RAM. Re-Price Your Server BOM.

Nvidia carries $279B in supply commitments, an increase it attributes primarily to memory, and still guided gross margin from 75% down to a 71-72% trough on RAM costs. Your server quote has no such contract behind it, and every self-hosted inference TCO model built on 2025 hardware prices is now understated.

August 27, 2026 · 11 min read
agent harness

Wiping the Transcript Was the Worst Retry. Paraphrase.

A new arXiv paper measures the corrective value of a failed tool call in an agent transcript and finds it negative on every small model tested: repeat probability rises from 0.06 to 0.54. Clean-context retry was the worst harness tested; paraphrasing the failure removed 76% of the effect for free.

August 26, 2026 · 13 min read
vLLM

vLLM vs TensorRT-LLM vs SGLang: Default to vLLM

TensorRT-LLM 1.2 removed the TensorRT engine build entirely, invalidating most published advice on this choice. vLLM is the default, SGLang wins prefix-heavy agent traffic, and the properties that actually matter never appear in a benchmark.

August 24, 2026 · 18 min read
MI355X

AMD's Inference Discount Depends on a GPU You Can't Rent

One of the first multi-vendor serving benchmarks on Kimi K3 puts AMD's MI355X at roughly 1.5x Nvidia B300's tokens per dollar — but that rests on a $2.50/GPU-hour rate you cannot currently buy. At the only purchasable on-demand rates, the gap narrows to 6%.

August 2, 2026 · 14 min read
self-hosted inference Articles | THE D*AI*LY BRIEF