OpenAI Decisions APIOpenAI's Decisions API Skips Output Fees and the Cache Discount
OpenAI's Decisions API bills only input tokens at $0.10 per 1M, which saves a little on short tickets. It also drops Luna's cached-input discount, so long shared rubrics can cost about four times more than the Responses API.
October 8, 2026 · 9 min readdecision modelsJev vs Clef vs Haiku: Decision Models Cut Cost 25x, Not Errors
Decision models such as Jev cost about 25x less per call than Claude Haiku 4.5 at list price, but trail LLMs on accuracy until you add labels and a calibration fit. Fine-tune if you have labels; cascade Jev into an LLM if you don't.
October 7, 2026 · 15 min readTypeSafe JevTypeSafe's Jev Scores 62.6% Asked Once and 95% Split Five Ways
TypeSafe's Jev answers classification calls for 12-27x less than Claude Haiku 4.5, but the first independent tests show its accuracy depends on how you split the question and its probabilities need recalibrating per question.
September 20, 2026 · 13 min readApple Foundation ModelsApple's On-Device Model Is 99% Sure. Sample It Five Times.
An independent audit of SystemLanguageModel.default — the ~3B on-device model Apple hands developers — found its self-reported confidence separates right from wrong at AUROC 0.47, below a coin flip, while it confabulated on 69.1% of false-premise questions and refused 18.4% of benign summarization requests. A k=5 consistency wrapper fixes it, at 28.2% coverage on factual QA.
August 25, 2026 · 15 min read