If your agents already call TypeSafe's Jev to route requests, pick tools and run guard checks, you can now run that step on open weights you own, but only some of those decisions should move. Cloudflare's new Clef beats Jev on intent classification and tool-call selection in Cloudflare's own benchmark, loses on the "should I act at all" questions, and costs about 5.7 times as much per input token on Cloudflare's hosted API. Amazon's Strands Decider 2B is free to run on a single consumer GPU and is weaker still on hard problems. Split your decisions by type, move the classification ones, and shadow-test before you cut over.
Both releases landed on October 1, 2026. They are open-weight decision models from two infrastructure vendors, in a class Jev opened in September.
What a Decision Model Is, and What Shipped on October 1
A decision model is a model that reads a state (text, JSON, sometimes images) and returns typed answers with probabilities instead of generating text. You hand it a schema of questions (yes or no, pick one of these options, score this from 1 to 5) and it hands back calibrated answers. The output is small and parseable, so it fits the parts of an agent that are really classifiers: which tool to call, which model to route to, whether a tool argument looks wrong, whether a ticket is in scope.
TypeSafe's Jev opened the category in September. It is a hosted API priced at $0.042 per million input tokens with output free, and the same docs list a 64k-token request limit with 32k tokens reserved for the state plus the longest question. Cloudflare resells it on its own platform at the same $0.042 input price, with zero data retention. We covered how Jev behaves in practice, including why splitting one question into several raises its accuracy sharply, in our September piece on Jev's calibration.
What changed this week is supply. Cloudflare released Clef and Clef-flash, the first models its Workers AI team trained itself. Clef is post-trained from Qwen 3.8-27B and includes a vision encoder; Clef-flash is a 9B sibling. Cloudflare calls both "fully Jev-API compatible" with strictly typed outputs, so the request your code already sends to Jev is the request Clef expects. The weights are on Hugging Face, where the Clef-flash model card states an Apache-2.0 license, inherited from its Qwen base, and says the model was tested on a single H200. Cloudflare also announced an RL fine-tuning service, delivered by forward-deployed engineers now and a self-serve platform later.
The same day, AWS's Strands Labs released Strands Decider 2B: a Qwen3.5-2B torso with its language-model head removed and replaced by a pointer head of just over a million parameters, plus a rank-16 LoRA adapter. Amazon published the weights, the training data and the training scripts. That last part matters, because Developers Digest's write-up of Clef points out that Cloudflare's training data and pipeline are unpublished, which makes Clef open weights rather than a reproducible open-source release.
Where Clef Beats Jev, on Cloudflare's Own Index
Clef wins clearly on the decisions that are really classification: mapping an utterance to an intent, or a request to a tool. The numbers come from the Jev Decision Index in Cloudflare's launch post, a set of 43 evals. Treat them as the vendor's own measurement, because they are; Cloudflare chose the benchmarks, ran them and published them in its launch post.
| Benchmark (Cloudflare's Jev Decision Index) | Clef | Clef-flash | Jev |
|---|---|---|---|
| BANKING77 intent, macro-F1 | 94.20 | 90.93 | 79.74 |
| CLINC150+OOS intent, macro-F1 | 97.43 | 66.77 | 89.27 |
| BFCL tool call, case exact | 98.47 | 98.76 | 95.75 |
| API-Bank, accuracy | 91.93 | 93.11 | 88.19 |
| Home appliances, case exact | 82.95 | 97.73 | 52.27 |
| When2Call, accuracy | 72.37 | 65.58 | 80.97 |
| BRIGHT retrieval, nDCG@10 | 45.91 | 39.26 | 47.52 |
| Agent trace observability (workflow eval) | 68.5 | 69.8 | 71.6 |
BANKING77 is the result a bank or fintech support team should look at first. A 14-point macro-F1 gap on 77 banking intents is the difference between a router you can leave alone and one that needs a fallback queue.
Read the Clef-flash column before you pick the cheap one. On CLINC150 with out-of-scope detection, Clef-flash scores 66.77, below Jev's 89.27 and 30 points under its larger sibling. Out-of-scope detection is the part of intent routing that keeps a request about a mortgage from landing in the card-dispute flow. If your router has to say "none of the above", the 9B model is the wrong default on this evidence.
Latency is the other win. Cloudflare reports a median of 38.8 ms for Clef-flash, 209.3 ms for Clef and 524.1 ms for Jev across its benchmarks. In a loop where an agent makes a dozen routing and guard decisions per task, a half-second per decision is visible to the user.
Where Jev Still Wins
Jev holds the lead on decisions that need judgment rather than lookup. On When2Call, which tests whether a model should call a tool, ask a follow-up or decline, Jev scores 80.97 to Clef's 72.37. Jev also edges ahead on BRIGHT, a reasoning-heavy retrieval benchmark, and on Cloudflare's agent-trace observability workflow eval.
The gap widens on general reasoning. Developers Digest reports Jev at 78.3 on GPQA Diamond against Clef's 48.0, and 82.7 on MMLU-Pro against 65.9. Those two figures do not appear in Cloudflare's post, so weigh them as secondary reporting, but they point the same way as When2Call.
Strands Decider sits further down that curve, and Amazon says so. Its launch post calls it "significantly worse at solving complex problems than reasoning models." MarkTechPost's breakdown of the JevBench v1 results shows 0.723 accuracy on the public set, with a perfect 1.000 on easy tasks, 0.875 on standard and 0.505 on hard ones. A coin flip on hard decisions is fine for picking which of four tools to call and a poor choice for approving a refund.
The pattern across all three models is consistent. Small open decision models are strong where the answer is in the input and the job is to recognise it. They are weaker where the answer requires weighing consequences, and that is where the expensive mistakes in an agent usually happen. One result in Cloudflare's table cuts the other way: on PhishNChips, a phishing-detection test, Clef scores 79.60 to Jev's 62.55, so a safety check that is really pattern recognition may belong on the classification side of your split.
What Hosted Clef Costs Next to Jev
On Cloudflare's hosted API, Clef costs more than Jev. Workers AI pricing lists Clef at $0.240 per million input tokens and Clef-flash at $0.090, with no output charge. Jev is $0.042 on TypeSafe's own docs and on Cloudflare's catalog. That makes hosted Clef about 5.7 times Jev's input price and Clef-flash about 2.1 times.
Run it at volume and the gap is easy to see. Developers Digest's example of 1 million decisions on a 2,000-token state works out to 2 billion input tokens: $480 on Clef, $180 on Clef-flash. The same volume on Jev at $0.042 per million is $84.
For most enterprises none of those numbers will decide anything. At decision-model prices, accuracy on your decision types, latency inside the agent loop, and where the state is allowed to go matter more than the bill.
That shapes the case for self-hosting. Beating Jev's API on cost is hard at $0.042 per million tokens. The stronger reasons are that the state you are classifying (a customer's account record, a security alert, an internal document) cannot leave your environment, or that you want to fine-tune on your own labels and keep the result.
When Self-Hosting Clef or Strands Decider Makes Sense
Self-host when data residency or fine-tuning is the requirement, and pick the model by the hardware you already have. Clef-flash was tested on a single H200, per its model card. Strands Decider runs on far less: Amazon reports a median of around 115 ms on an Nvidia RTX 3090 and around 153 ms on an M3 MacBook for small tasks. That makes Strands the option for an edge deployment, an air-gapped site or a developer laptop, and Clef the option for a GPU cluster you already run with vLLM or similar.
Strands also has a tighter fit if your agents are built on the Strands Agents SDK. The launch post shows it plugged into the SDK's intervention system as a check that runs before each tool call, which is the guard-check use case in a few lines of code.
Two things to settle before either goes to production:
- Authentication: MarkTechPost notes that the Strands Decider HTTP server binds to localhost with no authentication, so production use needs your own auth layer in front of it. A decision endpoint that approves tool calls is a control point and needs the same access controls as one.
- Logging: a local model call does not show up in your AI gateway's logs unless you route it there. We saw the same gap with Meta's laptop-sized agent model. If the decision gates an action, its input, output and probability need to land in your audit trail.
Hosted Jev keeps one advantage you lose when you self-host: someone else maintains the model. The Clef weights you download today are fixed. If Cloudflare ships a better Clef, you re-evaluate and redeploy, the same discipline we described for pinning and mirroring weights.
How to Decide, One Decision Type at a Time
Sort every decision your agents make into classification and judgment, then test each class separately. Classification decisions (intent, tool selection, model routing, category assignment) are where Clef's benchmark lead is large and consistent. Judgment decisions (should the agent act, is this request safe, does this output meet the policy) are where Jev and larger models still lead.
This is the same split we argued for in the model router buyer's guide: automate the mechanical routing and keep a stronger model on the call that carries risk. Cloudflare's own auto router results are a reminder that a vendor's routing layer can underperform on the vendor's own test.
This Week:
- Pull a week of Jev calls from your logs and tag each one by decision type: intent, tool choice, model route, guard check, policy judgment. You need the mix before you can price a move.
- For the intent and tool-choice set, sample 500 real calls with their outcomes and replay them against Clef on Workers AI. The API is compatible, so this is a base URL and model name change.
- If you route with an out-of-scope or "none of the above" option, include those cases on purpose. That is where Clef-flash fell to 66.77 on CLINC150.
This Month:
- Run Clef or Strands Decider in shadow mode beside Jev for two weeks on the classification decisions only: both answer, Jev's answer acts, and you compare agreement and accuracy against labelled outcomes.
- Leave guard checks and "should I act" decisions on Jev or a larger model until a shadow run on your own data shows otherwise. When2Call is the closest public proxy, and Jev leads it by 8.6 points.
- If residency is the driver, have your security lead sign off on the self-hosted endpoint's auth and logging before any traffic moves. Amazon's server ships without auth.
Before Renewal:
- If you have a Jev commitment, use the shadow results to size what share of volume could move. Even at Jev's low per-token price, a credible open-weight alternative for most of your classification traffic is leverage.
The Bottom Line
The open decision models that shipped on October 1 reopen the build-versus-buy question for one layer of the agent stack. A month ago the reference model in this class was a hosted API with proprietary weights. Now Cloudflare and Amazon both publish Apache 2.0 weights that speak its API, and Cloudflare's own numbers show its model winning on routing and losing on judgment.
That is the same shape as most open-versus-hosted choices in this stack: the open model catches up first on the narrow, well-labelled task, and the hosted one keeps the harder cases for a while longer. Your logs will tell you which of your decisions are which.
Start with the log export and the 500-call replay.
Continue Reading
- TypeSafe's Jev Scores 62.6% Asked Once and 95% Split Five Ways
- Model Router Buyer's Guide: Buy Failover, Not Judgment
- Cloudflare's Auto Router Fails 4x as Often as Opus on Its Own Test
- Best Air-Gapped LLM Stack: Apache Weights on vLLM Beat NVIDIA's Fee
- Meta's Agent Model Fits on a Laptop. Nothing Logs It.
- Wiping the Transcript Was the Worst Retry. Paraphrase.
