Microsoft Decision-1Microsoft Ships Decision-1 at Jev's Price, Ranked on Unnamed Tests
Microsoft's Decision-1 lists at $0.042 per million input tokens with free output, the same price as TypeSafe's Jev. Its 'highest accuracy' claim rests on 36 unnamed benchmarks with no published scores or calibration data.
October 10, 2026 · 10 min readRAG evaluationRAG Eval Dataset: 200 Hand-Written Questions Beat Synthetic Ones
Hand-write about 200 RAG eval questions from real traffic, with gold answers owned by one domain expert. Score retrieval and generation separately, gate every PR, and use synthetic generators only for retriever coverage.
October 4, 2026 · 13 min readTypeSafe JevTypeSafe's Jev Scores 62.6% Asked Once and 95% Split Five Ways
TypeSafe's Jev answers classification calls for 12-27x less than Claude Haiku 4.5, but the first independent tests show its accuracy depends on how you split the question and its probabilities need recalibrating per question.
September 20, 2026 · 13 min readHaize LabsBeacon Bought Haize Labs. Who Red-Teams Your Agents Now?
Beacon Software acquired AI red-teaming firm Haize Labs to run its internal AI platform, and the announcement says nothing about existing customers. Here is what to secure before renewal.
September 19, 2026 · 11 min readretrieval evaluationJPMorgan Ranked 62 Retrievers for $800. Pool the Judgments.
JPMorganChase compared 62 retrieval configurations on a production financial-news QA system for about $800, by judging the union of retrieved documents once and reusing 79.6% of the labels. Its top five embedding configurations spanned 0.007 MAP.
September 4, 2026 · 12 min readLLM evaluationBraintrust vs Langfuse vs Promptfoo: Don't Pay for the Gate
Promptfoo's CLI is the only one of the four that fails a build without you writing the enforcement code, and it costs nothing. The paid platforms sell what happens after the gate: stored traces, run-to-run diffs and a review queue.
August 26, 2026 · 17 min read