OpenAI's Decisions API Skips Output Fees and the Cache Discount

OpenAI's Decisions API bills only input tokens at $0.10 per 1M, which saves a little on short tickets. It also drops Luna's cached-input discount, so long shared rubrics can cost about four times more than the Responses API.

By Rajesh Beri·October 8, 2026·9 min read
Share:
A wooden mail-sorting rack in a back office with dozens of labeled pigeonholes, a single paper ticket suspended mid-air between two slots, and a brass postage meter on the desk beside it.

Illustration generated using AI

OpenAI's new Decisions API charges only for input tokens, so a short ticket classified on it costs a little less than the same call on the Responses API. The catch is that it also drops the cached-input discount: if every call carries a long shared policy or rubric, you can pay about four times more than you do today. Price your own traffic shape before you move a single queue, and set your thresholds on labeled data, because OpenAI publishes no calibration numbers and tells you to do exactly that.

The endpoint went into public beta this week. Per Zeli's October 6 write-up it is a dedicated route that "returns typed answers about 10x faster than the Responses API", and OpenAI's own developer guide says it expects general availability "in the coming weeks", with no date.

What Is OpenAI's Decisions API?

The Decisions API is an endpoint, POST /v1/decisions, that takes text or images plus a list of named questions and returns a typed answer for each one instead of free text. It runs on one model: the guide lists gpt-6-luna as "the only model currently available."

There are three question types. A predicate returns a probability from 0 to 1 that a condition is true. A choice returns one of the values you supply, plus a probability for each option and a separate confidence field. A score returns a probability-weighted average across ordered levels, which, as the guide notes, "can fall between levels." Several independent questions can share one request; a question that depends on another's answer needs a second call.

That puts it in the same category as the decision models we compared in our Jev, Clef and Haiku buyer's guide: a model that answers a closed question with a number you can threshold. What OpenAI adds is the distribution channel. Simon Willison has already published an LLM plugin for it, and on Hacker News he argued that an endpoint shape OpenAI defines tends to become the one other providers copy.

The announcement trail was messy. OrcaRouter's September 30 analysis described a limited preview announced at DevDay on September 29, with no published price. eesel reported on October 2 that /v1/decisions returned HTTP 403 for its account and that the guide URL was a 404 (eesel). The guide is live now and carries the price.

Where Does Input-Only Pricing Save Money?

It saves money on short items with no shared preamble, and it loses money when every call repeats a long instruction block. The guide states the rate: "$0.10 per 1M input tokens... You pay only for input tokens: there are no cache-read, cache-write, or output-token charges." That cuts both ways. On the same model through the Responses API, OpenAI's pricing page lists GPT-6 Luna at $0.10 per 1M input, $0.01 per 1M cached input, $0.125 per 1M cache writes and $0.50 per 1M output. On October 8 that page had no Decisions row at all.

Here is the arithmetic per million calls, at those list prices, for two common traffic shapes. It is our calculation, and the Responses figures assume a 10-token answer, reasoning off, and a warm cache on the shared prefix. TypeSafe Jev is included at its published $0.042 per 1M input, output free.

Per 1M calls Responses API, Luna Decisions API TypeSafe Jev
500-token ticket, no shared prompt $55 $50 $21
2,000-token cached rubric + 300-token item $55 $230 $96.60

On the first shape, Decisions saves about 9%. The output tokens you stop paying for were a small slice of the bill. The bigger saving shows up if you currently run Luna with reasoning on: eesel measured its own Luna triage at $0.047 per 1,000 tickets with reasoning off and $0.089 at medium effort, and output-side reasoning is exactly what the Decisions meter does not bill.

On the second shape, the one most moderation and policy-routing teams actually run, the rubric that cost $0.01 per 1M tokens as a cache hit costs $0.10 on every call. Decisions comes out at roughly four times the Responses bill. Two surcharges also carry over: the guide says "regional processing premiums and long-context input pricing multipliers apply", and the pricing page sets the regional uplift at 10% and doubles Luna's input rate above 272K tokens. If your cache hit rate is poor, the gap narrows; our piece on why one timestamp can turn 80% caching savings into 6% explains how to check yours.

How Far Can You Trust Its Probabilities?

Trust them for ranking, and calibrate them yourself before you threshold on them. The guide says "use labeled examples from your application to set thresholds for routing, filtering, or review" and to choose thresholds "based on the cost of false positives and false negatives." It does not say how the confidence field is computed or publish any calibration curve.

The first public tests point in two directions. On Hacker News, user armcat ran loaded-coin and marble-draw experiments and found that predicate questions gave "nearly perfect/expected probability outcomes", while a choice question turned a true 50% red-marble draw into 86%. Moving red from the first option to the last moved the estimate to 73%, on a question whose right answer was known in advance.

Another commenter, Topfi, ran fewer than 600 calls through OpenRouter against Jev and Mercury Decide, measured Decisions at 346ms p50 and 860ms p95, called it slower and about 3.1 times more expensive than Jev on average, and logged 4 failed calls against 0 for each competitor. A third commenter, brianyu8, who listed an openai.com contact address, re-ran two of Topfi's examples in the same thread and got 0.99 and 1.0 where Topfi had reported low scores. These are single-user tests with small samples, and they disagree with each other. Treat them as a reason to run your own, which is the conclusion we reached for Jev when its single-question accuracy jumped once the question was split five ways.

Latency has the same split. OrcaRouter reports OpenAI's own figure as "on the order of 150 milliseconds against about 1.6 seconds" for regular Luna, and calls it vendor-stated and unreproduced. One HN tester saw 160 to 175ms end to end with a worst case of 743ms; Topfi's p95 was 860ms. Neither number comes with a service level, because there isn't one in beta.

What Does the Beta Leave Out?

Most of the gaps are ordinary for a beta, and two of them matter for regulated data. Images "must be inline base64 data URLs"; hosted URLs and file_id inputs are not supported, per the guide, so an image pipeline that passes storage links has to fetch and encode first. The guide publishes no context window, no maximum number of questions or labels and no rate limits.

Data controls are present but conditional. OpenAI's data controls page marks /v1/decisions as Zero Data Retention eligible "with limitations" and its PSP and Safety Retention eligibility as "Pending confirmation." ZDR is "subject to prior approval by OpenAI", HIPAA use needs an executed Business Associate and Healthcare Addendum, and Europe regional processing requires ZDR or one of the modified retention arrangements. Residency covers the United States and Europe (EEA plus Switzerland). If your contract does not already include those approvals, the fast path is a sales conversation.


Should You Move Classification Traffic to It?

Move short, latency-sensitive calls you already run on Luna; leave long-rubric moderation and anything already on a cheaper decision model where it is. The strongest case for Decisions is operational: one vendor, one key, ZDR and HIPAA paperwork you may already have, US and EU residency, and a sub-second answer on a call that sits inside a user-facing turn. If your router or agent picks a tool on every turn, a few hundred milliseconds saved per decision can be worth more than the per-token price. Our model router buyer's guide covers the wider routing decision.

The case against is price and maturity. Jev lists at less than half the per-token rate, and Cloudflare's Clef beat Jev on routing. And the endpoint runs only on Luna, so whatever ceiling Luna has on hard judgment calls, the Decisions endpoint inherits; our Sol and Luna cost-per-task analysis covers where that tier falls short.

What to Do Before You Route Anything on It

This Week:

  1. Pull one week of the yes/no and pick-one calls you send to the Responses API and split them by shape: median input tokens, the share that is a repeated prefix, your cache hit rate, and output tokens including reasoning. Run the two-row table above on your numbers.
  2. Build a labeled set of at least a few hundred real items per decision, with the cost of a false positive and a false negative written next to each one. You need this regardless of vendor; the guide assumes you have it.

This Month:

  1. Shadow-run Decisions, your current Luna prompt and one decision model on that set. Compare cost per call, p95 latency, failed calls and accuracy at the threshold you would actually use.
  2. Where you use choice questions, re-run each item with the options shuffled. If the winning probability moves the way armcat's did, rewrite the question as separate predicates.
  3. Ask your OpenAI account team in writing whether your ZDR or Modified Abuse Monitoring approval covers /v1/decisions, and what "with limitations" means for your data.

Before GA:

  1. Get the Decisions price onto a quote or the public pricing page before you commit volume. A rate that lives only in a beta guide can change at GA.
  2. Keep your existing classifier path live behind a flag until the endpoint has a published rate limit and a service level.

The Bottom Line

OpenAI's endpoint turns classification into a metered utility on the same terms as Jev and Clef, and it brings the procurement advantages of a vendor most enterprises have already approved. The pricing model rewards short, unique inputs and penalizes the long, repeated rubrics that moderation and policy routing depend on. Teams that saw cheap classifiers arrive with Jev already have the playbook: measure cost per decision on your own traffic, then calibrate on your own labels.

Run the two-row cost table on last week's logs before anyone opens a migration ticket.

Continue Reading

Share:

Frequently Asked Questions

What is the OpenAI Decisions API?

It is a beta endpoint, POST /v1/decisions, that takes text or images plus named questions and returns typed answers: a probability for a predicate, one option with per-option probabilities for a choice, or a weighted score across ordered levels. It runs only on gpt-6-luna.

How much does the OpenAI Decisions API cost?

OpenAI's developer guide lists $0.10 per 1M input tokens, with no output, cache-read or cache-write charges. Regional processing premiums and long-context multipliers still apply. As of October 8, 2026 the main API pricing page had no Decisions row.

Is the Decisions API cheaper than calling GPT-6 Luna through the Responses API?

For short items with no shared prompt it is about 9% cheaper, because you stop paying for output tokens. If every call repeats a long cached rubric, it can cost roughly four times more, because Luna's $0.01 per 1M cached-input rate does not exist on the Decisions endpoint.

Are the Decisions API probabilities calibrated?

OpenAI publishes no calibration data and tells developers to set thresholds on labeled examples from their own application. Early public tests found predicate questions well behaved, while a choice question's probabilities shifted when the order of options changed.

Does the Decisions API support zero data retention and HIPAA?

OpenAI marks /v1/decisions as ZDR eligible with limitations, subject to prior approval, and HIPAA eligible under an executed Business Associate and Healthcare Addendum. Data residency covers the United States and Europe (EEA plus Switzerland).

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →