Topic

AI inference cost

Every THE D[AI]LY BRIEF article on AI inference cost — enterprise AI analysis, benchmarks, vendor comparisons, and ROI frameworks for technology and business leaders. Updated as new coverage publishes.

KV cache

Random Eviction Matched the Scorers. Protect the Prompt.

Salesforce AI Research and UIUC replaced the scoring pass in KV cache eviction with a uniform random draw, kept the prompt pinned, and matched the strongest published evictor across four models and six reasoning tasks — at 32-43% higher vLLM throughput. Random-plus-prompt-protection is now the baseline any 'intelligent cache compression' claim has to beat.

September 9, 2026 · 13 min read
AI capacity procurement

AWS Gets Paid to Buy. Don't Lock Your Instance Family.

Qualcomm's press release called it a collaboration; its 8-K disclosed a warrant on 25,000,000 shares at $161.26 that vests as Amazon places binding purchase orders, up to $60 billion, through 2036. The silicon is already in production and AWS has announced no instance type, which makes contract language the only lever you hold this quarter.

September 8, 2026 · 15 min read
prompt caching

Anthropic Cut Cache Reads 75%. Opus 5 Still Undercuts It.

Anthropic cut Claude Fable 5.1 cache reads 75% to $0.25/MTok, half of Opus 5's rate on a model with twice the base input price. But Opus 5 stays cheaper until cache reads pass a third of your bill, and Anthropic's own savings ceiling of 45% falls short of the 50% crossover.

September 1, 2026 · 15 min read
Claudeforce

Claudeforce Runs Two Meters. One Throttles, One Bills.

Claudeforce puts 37 prebuilt sales skills inside Claude, with open beta in September 2026 and no published price. Salesforce meters MCP tool calls against your org's shared daily API allocation; Anthropic bills inference on a separate contract with no ceiling.

August 28, 2026 · 14 min read
GPT-5.6 Sol

GPT-5.6 Sol Is $20 Until Nov 21. Budget Both Rates.

OpenAI cut GPT-5.6 Sol to $4/$20 per million tokens but guarantees the rate only "at least through November 21, 2026," with no successor published. The cut was asymmetric, so no two workloads saved the same amount — and the reversion is +50% on output, not +33%.

August 24, 2026 · 12 min read
Decart

Anthropic Walked From Its $6B Decart Bid. Don't Fix Your Rate.

Anthropic is in advanced talks to buy Decart for roughly $7 billion, mostly in its own pre-IPO stock, to cut what a Claude token costs it to produce. Every Sonnet through 4.6 still lists at its March 2024 price, and the one cut that did land arrived as a new model number — which is why a flat multi-year rate card is the wrong thing to sign this quarter.

August 18, 2026 · 13 min read
DeepSeek V4

DeepSeek Will Raise Prices. Your Ceiling Is Already 4x.

DeepSeek warned of a significant API price rise with no number and no date. Because the weights are MIT-licensed, the ceiling is already public: independent hosts serving the identical model charge 3-4x on posted rates and 40x on cache hits.

August 7, 2026 · 12 min read
Snowflake Cortex

Snowflake Cortex vs Databricks Mosaic AI: Pick on Exit Cost

Databricks lists 47 servable models to Snowflake's ~30, but Unity Catalog keeps Delta tables read-only to outside engines and Snowflake's own catalog blocks third-party writes outright. Snowflake's exit advantage is narrower than the decks claim: it writes into a catalog you own. On a normalised 10-million-ticket workload the platforms land within 2x of each other; the model tier swings the bill 40x.

August 3, 2026 · 16 min read