token inflationA 40-Call Audit Flagged 7 of 15 GPT Endpoints for Token Inflation
A 40-call black-box audit from USTC flagged 7 of 15 OpenAI-compatible GPT services for behaviour consistent with output-token inflation. The flags are not proof, but padded answers pass quality evals and invoice checks alike, so measure output tokens per task against the first-party API.
September 21, 2026 · 13 min readprompt injectionAn Eval Sandbox Gave Up Its Keys. Your Gateway Holds Yours.
Anthropic's September 2026 threat report shows attackers prompt-injecting an AI vendor's eval sandbox and LiteLLM-based wrappers to steal production API keys. Any harness or gateway that reads untrusted text while holding a key is a credential store — split, scope and cap those keys.
September 12, 2026 · 11 min readmodel routingModel Router Buyer's Guide: Buy Failover, Not Judgment
A model router sells two things: mechanical failover, which works, and a learned classifier that picks your model, which the neutral benchmarks say does not. Buy the first, build the second, and plan on 35% savings rather than 60%.
September 11, 2026 · 20 min readLLM gatewayBest LLM Gateways for Cost Control: Self-Host First
Self-host LiteLLM: per-team budgets and virtual keys are in the free open-source tier, while everyone else gates enforcement behind a sales call. Priced through one 50M-request workload, the platform layer ranges from $7 to $10,300 a month.
August 7, 2026 · 19 min read