If an LLM call returns one word from a fixed list that no human reads, a decision model will do that job for about a twenty-fifth of what Claude Haiku 4.5 costs. It will not do it as accurately out of the box. The best independent study so far found decision models a median 11.6 macro-F1 points behind the best LLM on each task, at 44 times lower cost.
So the recommendation splits on one question: do you have labels? If you have 1,000 or more labelled examples and criteria that hold still, fine-tune a small classifier and own it. If you don't, put TypeSafe's Jev in front of your existing LLM as a first pass and send its low-confidence answers to the LLM. Self-host Cloudflare's Clef or Amazon's Strands Decider only when the data cannot leave your network. Stop sending every decision to an LLM with a JSON schema bolted on.
| Option | What it returns | Price (checked Oct 7, 2026) | Best at | Do NOT pick it if |
|---|---|---|---|---|
| TypeSafe Jev (hosted) | Typed Choice / Score / yes-no with probabilities | $0.042 per M input tokens, output free | Changing criteria, no labels, judgment-style calls | You need images, non-English, or a regulated path today |
| Cloudflare Clef / Clef-flash (open weights) | Jev-compatible typed answers | Hosted: $0.240 / $0.090 per M input tokens; self-host: your GPUs | Intent and tool routing, image input | You need "should I act at all" judgment, or will pay hosted prices |
| Strands Decider 2B (open weights) | Yes-no or option pick with confidence | Free weights; runs on one consumer GPU | Edge, laptop, air-gapped guard checks | The decision is hard; it scored 0.505 on hard tasks |
| Fine-tuned small classifier (ModernBERT, LoRA on a 4B model) | Class label with a probability you calibrate | Your training time plus serving | Stable criteria with 1,000+ labels | Your categories change every quarter |
| Claude Haiku 4.5 with structured outputs | Schema-valid JSON text | $1 per M input, $5 per M output | Low volume, reasoning-heavy calls, the fallback tier | It runs on every decision at high volume |
What Is a Typed Decision Model, and How Is It Different From JSON Mode?
A typed decision model reads a state and returns one of the answers you defined, with a probability attached, and it never generates text. You send TypeSafe's Jev a state (a string, a JSON object or an array of text) plus typed questions, and it answers each one in a single pass. TypeSafe's docs list jev-1.13.0 at $0.042 per million input tokens with output tokens free, a 64k-token request budget with 32k reserved for the state plus the longest question, and text-only input.
An LLM with structured outputs looks similar from the outside and works differently inside. Anthropic's structured outputs constrain decoding so the response is "always valid" against your schema, with "no retries needed for schema violations". OpenAI's version makes the same promise: the model will not omit a required key or hallucinate an invalid enum value. What you get back is still generated text that happens to parse. You pay output tokens for it, and any confidence number you ask for is the model writing a number down.
A parsed LLM string gives you a label. A decision model gives you a probability distribution over labels, which is what lets software decide on its own when to act and when to escalate.
Is Calibrated Confidence Worth Paying For?
Yes, and it's the main reason to move. Confidence you can threshold is what turns a classifier into automation, because it tells you which share of traffic needs no human. A model is calibrated when answers it gives at 0.9 probability are right about 90% of the time.
The evidence says decision-model confidence beats what LLMs write about themselves, and still needs fitting. Ibrahim and Zaki's study of 18 social-science classification tasks (7,977 items, 19 LLMs) found decision models better calibrated than the verbalised confidence of 16 of the 19 LLMs. Three frontier models still did better, with a median calibration error as low as 0.066 against Jev's 0.157. Above 0.9 confidence, decision-model answers were right at a median of 0.815, except on empathy tasks, where they were near chance.
A crash-records study makes the same point at production scale. Rafe and Das used Jev to code 195,857 Texas police crash narratives against a 27-question schema. Jev scored F1 0.908 against human labels, one frontier LLM beat it by 0.059 and the other was statistically tied, and recalibration cut calibration error by a factor of 3.3. Their conclusion is the one to plan around: calibration varies model by model, so each one needs its own validation.
Out of the box, Jev's probabilities drift by question type. An independent calibration study on 900 synthetic support tickets found yes-no answers underconfident and Choice and Score answers overconfident, with an expected calibration error of 0.107. Budget for a per-question calibration fit on held-out labels. That is a few hundred labels per question, which you need anyway to know whether any option works.
Measured Cost and Latency Against an LLM Baseline
At list prices on a defined workload, Jev costs about 25 times less than Haiku 4.5. Here is the workload, so the comparison is like for like: 1 million decisions a day, each with a 1,500-token state and question set, one five-way Choice, about 20 output tokens for the LLM.
| Option | Daily cost at list price | Source of the rate |
|---|---|---|
| Jev | 1.5B input tokens × $0.042/M = $63 | TypeSafe models page |
| Clef-flash (hosted) | 1.5B × $0.090/M = $135 | Workers AI pricing |
| Clef (hosted) | 1.5B × $0.240/M = $360 | Workers AI pricing |
| Haiku 4.5, structured output | 1.5B × $1/M + 20M × $5/M = $1,600 | Anthropic pricing |
| Haiku 4.5, Batch API | half of the above = $800, asynchronous only | Anthropic pricing |
Self-hosted Clef, Strands Decider and a fine-tuned classifier cost whatever your GPU hours cost, so they are not priced here. All prices were checked on each vendor's live page on October 7, 2026.
The independent measurement agrees with that arithmetic. A reproducible phishing benchmark on 2,000 emails put Jev at $0.038 per 1,000 emails against $0.462 to $1.02 for Haiku, 12 to 27 times cheaper depending on how many questions were asked. Median latency was 239 ms for Jev against 687 ms for Haiku, measured from France. InfoQ's launch report carries the user-reported numbers: a median of 76 ms, a safety classifier running 5 to 18 times faster than the LLM it replaced (reported by a Vercel engineer), and a 30x median cost saving across launch-week posts. Treat that last figure as anecdote. Nobody controlled the baselines.
Cloudflare's Clef launch post reports median latency of 38.8 ms for Clef-flash, 209.3 ms for Clef and 524.1 ms for Jev on its own benchmark. Those are a vendor's numbers on a vendor's index.
One constraint matters before the bill does. TypeSafe now lists rate limits of 100K tokens a second and 80 requests a second, and says they are "adjusting dynamically". At 1,500 tokens a request, 80 requests a second is about 6.9 million decisions a day at a perfectly flat rate. Real traffic peaks. If your busiest hour runs at five times the average, a 1-million-a-day workload is close to that ceiling, so ask for a committed limit in writing.
Does a Schema Remove the Failure Mode or Hide It?
A schema removes one failure, output your code can't parse, and makes the costlier one harder to see: a valid answer that is wrong. Jev cannot return a label outside your enum, and neither can Haiku with strict structured outputs. Both can return the wrong label with full confidence.
Several failures get easier to miss once the parse errors disappear:
- Wrong answers that parse. TypeSafe's own jaggedness page lists counting, date comparisons, double negatives and noisy state as unreliable. It also says Jev "leans toward the option that comes first" in a Choice. A parsed string was at least visible when it was garbage. A typed wrong answer looks the same as a typed right one.
- No "none of the above". A pre-registered evaluation of Jev found 0 of 30 out-of-scope messages flagged, at 0.99 confidence. If your Choice has no escape option, the model picks the least wrong one. Add it to every Choice.
- The question text becomes the program. The same pre-registered study found that wrong criteria descriptions scored 16.7%, below the 25% random floor. Review question wording the way you review code.
- Refusals and truncation on the LLM side. Anthropic's docs say a refusal or a
max_tokensstop can return output that does not match your schema, with a 200 status code. Your "guaranteed" parser still needs a branch for that. - Format constraints can cost accuracy. Tam et al. found "a significant decline in LLMs reasoning abilities under format restrictions". If your LLM call does real reasoning, forcing a one-token JSON answer can make it worse at the job.
- Prompt injection. TypeSafe says content "written to adversarially steer the model... can move the answer." A guard check that reads attacker-written text needs a dedicated injection detector in front of it.
Each Option, and Who Should Not Buy It
TypeSafe Jev is the default first pass for teams without labels. The case for it is agility: you change the task by editing the question text, with no retraining. Its accuracy depends on how you frame the question. In the phishing benchmark, Jev scored 62.6% asked once and 95.0% when the task was split into five narrow questions with weights fitted on 1,000 labels. Skip Jev if your state is mostly non-English (TypeSafe says English is where accuracy is best), if you need image input, or if the decision is regulated and you can't yet get zero data retention and a security attestation in a contract. The models page offers a DPA and enterprise ZDR and lists no SOC 2 or ISO report. We went deeper on its quirks in our September analysis of Jev.
Cloudflare Clef is the open-weight option for routing. Cloudflare's index has Clef at 94.20 macro-F1 on BANKING77 intents against Jev's 79.74, and behind Jev on When2Call (72.37 against 80.97), the test of whether to act at all. The Clef-flash model card lists an Apache-2.0 licence, and Clef carries a vision encoder Jev lacks. Hosted Clef is the loser in this guide. It costs 5.7 times Jev per input token for a model that trails Jev on judgment. Download the weights and run them yourself if residency demands it, or pay for it hosted only if you need image input. Our Clef and Jev head-to-head has the full benchmark table.
Strands Decider 2B is for decisions that have to run on hardware you already own. Amazon's launch post reports a median of about 115 ms on an RTX 3090 and about 153 ms on an M3 MacBook, and says it is "significantly worse at solving complex problems than reasoning models." MarkTechPost's breakdown shows 0.505 accuracy on hard JevBench tasks, and the project README documents no authentication for its HTTP server. Don't use it for anything you would hesitate to hand a coin flip, and put your own auth in front of it.
A fine-tuned small classifier wins whenever you already have the labels. On the same phishing set, a commenter trained Qwen3-4B with LoRA on 1,000 emails and reached 97.4% accuracy with a calibration error of 0.010, in 18 minutes on one RTX 4070 SUPER. That beat every hosted option on both accuracy and calibration. An encoder such as ModernBERT (149M parameters base, 8,192-token context) is the lighter version of the same move. Don't pick this route if your categories change monthly, if no one on the team will own retraining, or if serving the model means standing up GPUs you don't otherwise run.
Haiku 4.5 with structured outputs belongs in the fallback tier. It still wins on reasoning-heavy calls and low volume, where $1,600 a day doesn't apply. Anyone running it on every routing and triage call at seven-figure daily volume should not keep doing so, and that includes larger models on the same job.
Guardrail and Routing Use Cases That Fit
Decision models fit the calls in an agent that are really classifiers: which tool to call, which queue a ticket goes to, which model handles a request, whether a tool argument looks wrong. Routing and intent are where open models already match or beat Jev on Cloudflare's numbers. Judgment calls such as "is this action safe" or "should the agent ask first" are where Jev and larger LLMs still lead.
The cascade is what makes the economics work without giving up accuracy. Ibrahim and Zaki found that routing low-confidence items to an LLM "matches or exceeds the LLM alone at a quarter to half of its cost." Your saving is the share of traffic that clears your confidence threshold, multiplied by the price gap. On the pre-registered study's data, the 0.99 band covered 60.2% of traffic at 100% accuracy. Your share will differ, so measure it.
Two guard-check cautions. A decision model that approves tool calls is a control point, so its input, output and probability belong in your audit log. And because it trusts the state by default, it should never be the only gate on text an outsider can write. For picking which LLM answers rather than replacing the call, see our model router buyer's guide.
The Limits Vendors Acknowledge at Launch
Every vendor in this class published its own caveats, and they should go straight into your risk register:
- TypeSafe: unreliable counting, arithmetic and date comparison; accuracy falls as irrelevant state grows; adversarial content can move answers; first-option bias in Choices; rate limits "adjusting dynamically"; English-first. Its models page also says Jev is not trained on customer requests or responses.
- Cloudflare: fine-tuning "may give up some general purpose performance in exchange for higher accuracy in a specific domain," per its launch post. The training data and pipeline are unpublished.
- Amazon: worse than reasoning models on complex problems, and the model "always produce[s] an answer from the selected options," so you supply the escape option.
- Anthropic and OpenAI: schema guarantees break on refusals and token-limit stops, and the first request with a new schema is slower while the grammar compiles.
Arize's review said it would run its own benchmarks before taking the vendor's word. Do the same.
How to Decide: The Criteria That Predict Regret
Three questions settle it. Answer them in this order.
- Do you have 1,000+ labels and stable criteria? Fine-tune. Every independent result in this guide favours the trained classifier on accuracy and calibration once labels exist.
- No labels, or criteria that change often? Jev as a first pass, an LLM on everything below your threshold, a shadow run before cutover.
- Can the state leave your network? If not, Clef on GPUs you run, or Strands Decider on smaller hardware, for routing and intent only. Keep judgment calls on a larger model you can host.
What changes the answer: a vision requirement (Clef), a non-English state (test hard before choosing Jev), a volume low enough that Haiku's bill is a rounding error (stay put), or a regulated decision with no vendor attestation yet (keep it on the LLM you've already contracted).
This Week:
- Export a week of LLM calls and tag every one whose output is a single label, score or yes-no that no human reads. That list is your addressable volume, and the price table above turns it into dollars.
- Pick the highest-volume one with ground truth, and pull 1,000 to 2,000 records where the label is the final human disposition.
This Month:
- Replay that set through Jev, your current LLM and, if labels allow, a LoRA or ModernBERT classifier. Add a "none of these" option to every Choice. Compare the cost of each correct decision.
- Fit calibration per question on half the set, score the other half, and find the confidence band where accuracy clears your bar. The share of traffic in that band is your saving.
Before Q1 Planning:
- Ship the cascade on one decision: decision model first, LLM below threshold, everything logged. Get a committed rate limit, zero data retention and a version pin in writing before the second decision moves.
The Bottom Line
This is the BERT-era classifier coming back, with one improvement: you can change the task by rewriting a question instead of retraining. The price gap is real and large. The accuracy gap is also real, and it closes only with labels, question design and a calibration fit you maintain.
Run the 2,000-record replay before you sign anything.
Continue Reading
- TypeSafe's Jev Scores 62.6% Asked Once and 95% Split Five Ways
- Cloudflare's Clef Beats Jev at Routing and Loses at Judgment
- Model Router Buyer's Guide: Buy Failover, Not Judgment
- Fine-Tuning vs RAG Cost: The Training Bill Isn't the Bill
- Only 3 of 36 Model Gaps Were Real. Size Your Eval Set.
- NeMo Guardrails vs Guardrails AI vs Lakera: Buy the Detector
- Braintrust vs Langfuse vs Promptfoo: Don't Pay for the Gate
