Microsoft Ships Decision-1 at Jev's Price, Ranked on Unnamed Tests

Microsoft's Decision-1 lists at $0.042 per million input tokens with free output, the same price as TypeSafe's Jev. Its 'highest accuracy' claim rests on 36 unnamed benchmarks with no published scores or calibration data.

By Rajesh Beri·October 10, 2026·10 min read
Share:
A server rack in a data centre aisle with a single physical rotary switch mounted on its front panel, its pointer resting between two blank positions, lit by cold overhead light. No text or logos.

Illustration generated using AI

If you run an LLM router, guardrail check or LLM-as-judge step, Microsoft now sells a decision model inside your existing Azure tenant at the exact price of the startup model you may be piloting. Microsoft-Decision-1 costs $0.042 per million input tokens with free output, the same list price TypeSafe publishes for Jev. Price no longer separates the two. What should decide it is accuracy on your traffic, and Microsoft's announcement gives you a ranking without the numbers behind it.

Microsoft published the model on October 9 through its Office of the CTO. It is post-trained from Alibaba's open-weight Qwen3.5-9B and available on Microsoft Foundry and OpenRouter, per the announcement. The post claims the "highest accuracy" across 36 benchmarks and "nearly 150,000 questions." It does not name the benchmarks or print a single accuracy percentage.

What Microsoft Actually Shipped

Decision-1 is a decision model: it reads a piece of text or JSON, answers a bounded question about it, and returns a probability instead of prose. The Foundry documentation describes three question types. A noul question returns a yes/no probability, choice picks one label from a fixed set with a probability for each, and score places an item on an ordered scale. Several questions about the same input can share one request.

The API shape will look familiar if you tested TypeSafe Jev. Both take a state plus named, typed questions, and the Foundry docs say the request pattern is "similar to a system-one API." The endpoint path on a Foundry resource is literally /providers/microsoft/v1/systemone. TypeSafe describes Jev as its "first System One model" on its models page. The request formats look alike, so test whether an existing Jev integration ports over before you budget for a rewrite.

Deployment options matter for regulated data. The docs list DataZoneStandard "in selected regions" and GlobalStandard. A Global deployment can process your inference in any Azure region where the model runs; a Data Zone deployment keeps processing inside the US or EU zone, per Microsoft's data, privacy and security page for models sold by Azure. Microsoft does not list which regions have the Data Zone option for Decision-1.

How the Price Compares

At list price, Decision-1 ties Jev, undercuts OpenAI's Decisions API and loses to Cloudflare's smallest Clef. Every option in this table charges for input only, as of October 11, 2026.

Model Input, per 1M tokens Output Source
Microsoft-Decision-1 $0.042 free Microsoft
TypeSafe Jev 1.13 $0.042 free TypeSafe
OpenAI Decisions API (GPT-6 Luna) $0.10 free OpenAI
Cloudflare Clef $0.240 none listed Cloudflare
Cloudflare Clef-flash $0.038 none listed Cloudflare

Run the volume and the spread is small at the cheap end. One million decisions on a 2,000-token state is 2 billion input tokens. That works out to $84 on Decision-1 or Jev, $76 on Clef-flash, $200 on OpenAI's Decisions API and $480 on full Clef (our arithmetic, at the list prices above). Clef-flash's price has dropped since we compared Clef and Jev on October 3, when Cloudflare listed it at $0.090.

A gap of $8 per million decisions will not move a procurement decision. The contracts you already hold will. If your Azure agreement covers Foundry consumption, Decision-1 arrives through an existing vendor relationship, while a startup API means a new vendor review, a new DPA and a new security questionnaire. At $84 per million decisions, that paperwork can easily cost more than the tokens.

Two pricing details are missing. OpenAI's guide says regional processing premiums and long-context multipliers apply to its Decisions API without listing them, which we covered in our Decisions API breakdown. Microsoft's post says nothing about whether Data Zone deployments cost more than Global for Decision-1. Ask before you model a regulated workload.


Why the Benchmark Claim Is Thin

Microsoft's accuracy claim is a ranking on tests you cannot inspect. The post says Decision-1 was evaluated on "dozens of benchmarks kept blinded from training, spanning routing, ranking, long context, multilingual and out-of-distribution tasks, reasoning, and safety," and tested on the JevBench leaderboard "across 36 additional public and private benchmarks." It never says which are public, which are private, or what anyone scored.

The comparison set has gaps of its own. The post's appendix lists six models it benchmarked: H2O-Lightning-4B, Quyet-1.0-Large, Surogate Rune 26B-A4B, deck-31B, Strands-Decider 2B and GPT-6 Luna Decisions. Jev is absent from that list, even though an editor's note says the post was updated "to add benchmarks for Jev on accuracy and calibration," and no Jev figure appears in the text. Cloudflare's open-weight Clef is not mentioned at all.

The speed numbers do not agree between sources either. Microsoft's post calls Decision-1 "2.5 times quicker than H2O-Lightning-4B v1.1, the runner-up, and 35 times quicker than GPT-6 Sol." The AI Weekly alert that carried the news reported it as 4.5 times faster than Quyet-1.0-Large. One of those may come from an earlier version of the post. Either way, the speed baseline that matters to you is your current router, and GPT-6 Sol is a frontier model nobody should be using to pick a ticket queue.

The internal case studies are also vendor claims. Microsoft says Xbox Research ran more than 10,000 feedback items at quality "competitive with GPT-6 Sol" while running "over 14 times faster" and "200 times less expensive," and that a Copilot team found it competitive with GPT-5.6 Luna for quality control. Those are useful as a signal of what Microsoft uses it for. They are not evidence about your labels. We wrote a guide to vendor benchmark claims on October 7 with a simple rule: only numbers with logs belong in the deck.

To be fair to Microsoft, the post does publish one measurable robustness result. Decision-1 changed its answer on 1.3% of perturbations on average across eight perturbation types, with zero flips when options were paraphrased, reversed or shuffled. That last part is the property you want in a router, and it is testable in an afternoon.

Calibration: The Docs Are More Careful Than the Blog

The announcement sells calibration; the documentation tells you not to rely on it. A calibrated model is one whose confidence matches its hit rate, so its 90% answers are right about nine times in ten. The blog post states that as the goal and says the probability is "part of the API, not just a ranking score." It reports no calibration measurement, such as an expected calibration error or a reliability curve.

The Foundry docs are blunter. For score questions they say: "Prefer scores for relative ordering and thresholds rather than as absolute, calibrated ratings." The limitations section says "Calibration is strongest on familiar task types," that "Scores can change based on how you phrase or order questions and options," and that "Poorly framed questions still return scores." Microsoft also tells you to validate thresholds "against labeled examples before you automate a gate."

That advice is correct, and it matches what OpenAI's guide says about thresholds for its own Decisions API. It also means the calibration claim cannot be the reason to switch. In September, TypeSafe's own Jev scored 62.6% asked once and 95% split five ways on the same task, which shows how much these models depend on how you frame the question. Expect the same of Decision-1 until you have measured it.

The Questions Your Security and Procurement Teams Will Ask

The base model is Chinese-origin, open-weight and Apache 2.0 licensed; Decision-1 itself has no published license. The Qwen3.5-9B model card lists Apache 2.0, and the announcement says Decision-1 is post-trained from it. If your organization restricts models by country of origin, as some public-sector and defense buyers do, get a ruling before the pilot starts. We covered how enterprises came to depend on Chinese base models in July.

The model under the API will change. Microsoft says it will rebase Decision-1 "on other models, including Microsoft AI (MAI) and OpenAI." The Foundry docs list the current deployment as model version "1." Pin that version, and treat a rebase as a new model that needs its thresholds re-validated, because a different base model can shift every probability you tuned against.

Data handling follows Azure's general terms, as far as Microsoft has said. Microsoft's data privacy page says prompts and completions for models sold by Azure are not used to train or improve the base models. It also says flagged content can be stored for human abuse review unless you are approved for modified abuse monitoring. Neither the announcement nor the Decision-1 docs say outright which data terms apply to this model, and the announcement says only that it runs in "a secure, trusted environment." If you route content that includes customer records through a decision gate, get that confirmed in writing.


What to Do Before You Switch Routers

Decision-1 is worth a pilot for any team already on Foundry. It has not earned a migration yet. Run it in shadow beside your current gate and let your labels decide.

This Week:

  1. Pull 300 to 500 labeled decisions from production logs for your highest-volume gate (routing, refusal check or rubric grade), including the borderline cases your team argues about.
  2. Deploy Decision-1 in a Foundry project on DataZoneStandard if your region offers it, pin model version 1, and replay that set through it, Jev and whatever you run now.
  3. Shuffle option order and paraphrase option text on a 50-case slice. Microsoft claims zero flips there; check it.

This Month:

  1. Bucket each model's probabilities into deciles and compare the predicted rate to the actual hit rate in each bucket. That is your own calibration curve, and it is the number Microsoft did not publish.
  2. Set thresholds per gate from the cost of a false positive versus a false negative, and add a "cannot tell" option where a forced choice would be wrong, as the Foundry docs suggest.
  3. Ask your Microsoft account team for three things in writing: the Data Zone regions for Decision-1, whether it is covered by the "models sold by Azure" data terms, and how much notice you get before a rebase.

Before Renewal:

  1. If Decision-1 matches your incumbent on your labels, price the consolidation: one fewer vendor review against the switching work. If it does not, you have a measured reason to stay and a benchmark to hold the next vendor to.

The Bottom Line

Since we covered Jev on September 21, the field has grown to five priced decision-model options: Jev, two Clef sizes, OpenAI's Decisions API and now Microsoft's. Three of them list at roughly four cents per million input tokens. Load balancers went the same way once every cloud shipped one, and the buying decision moved from the product to the integration and the contract.

Decision models are not at that point yet, because the outputs still differ underneath the matching prices. Microsoft has published neither accuracy nor calibration figures for Decision-1, and OpenAI's guide does not say how its confidence field is computed. Foundry makes Decision-1 easy to try, and the benchmark ranking tells you little about how it will do on your gates.

Replay your own labels through Decision-1 before you let it gate anything.

Continue Reading

Share:

Frequently Asked Questions

How much does Microsoft Decision-1 cost?

Microsoft lists Decision-1 at $0.042 per million input tokens, and output tokens are free. That is the same list price TypeSafe publishes for Jev 1.13, below OpenAI's Decisions API at $0.10 per million input tokens, and slightly above Cloudflare's Clef-flash at $0.038 (prices checked October 11, 2026).

What is Microsoft Decision-1 built on?

Microsoft says Decision-1 is post-trained from Alibaba's open-weight Qwen3.5-9B, which is Apache 2.0 licensed. Microsoft also says it will rebase the model on other models, including Microsoft AI (MAI) and OpenAI models, so pin the deployed model version and re-validate thresholds after any rebase.

Are Decision-1's probabilities calibrated?

Microsoft's blog calls them calibrated but publishes no calibration measurement. The Foundry docs say calibration is strongest on familiar task types, that scores change with question wording and option order, and recommend using scores for ordering and thresholds rather than as absolute calibrated ratings. Validate on your own labeled data.

Where can I deploy Microsoft Decision-1?

Decision-1 is available on Microsoft Foundry, with DataZoneStandard deployments in selected regions and GlobalStandard deployments, and Microsoft says it is also on OpenRouter. A Global deployment can process data in any Azure region where the model runs; a Data Zone deployment keeps processing inside the US or EU zone.

Should I switch from Jev or OpenAI's Decisions API to Decision-1?

Not on the published evidence. Microsoft's benchmarks are unnamed and carry no scores. Replay a few hundred labeled production decisions through Decision-1 and your current model in shadow, compare accuracy and calibration by decile, and switch only if it matches on your own traffic.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →