OpenAI's Ultrafast tier makes GPT-6 Astra generate tokens faster at six times the Standard price, and it only runs on US data residency or global processing. In the API that is $60 per million input tokens and $300 per million output tokens. For an Enterprise buyer that means three decisions: who gets it (it is off by default, and an owner switches it on for selected users or the whole workspace), what it burns (8x the Standard rate against included Codex limits), and whether your EU workloads can touch it at all (they can't). OpenAI's own documentation says the speedup applies to token generation and makes no claim about total task time.
OpenAI announced Ultrafast at DevDay on September 29, 2026, alongside a $500 Pro plan and cuts to the $200 one. Because it ships off for Enterprise workspaces, nothing has changed in yours unless an owner acted. Keep it off for everyone who doesn't have a measured reason to switch it on.
What Does Ultrafast Cost Per Million Tokens?
Ultrafast costs exactly six times GPT-6 Astra's Standard rate on every token line. VentureBeat's DevDay report puts Standard Astra at $10 input, $1 cached input and $50 output per million tokens, and Ultrafast at $60 and $300. Mixed's breakdown adds that cached input goes from $1 to $6, cache writes from $12.50 to $75, and that the 6x holds above the 272,000-token long-context threshold too.
| GPT-6 Astra, per million tokens (short context) | Standard | Ultrafast |
|---|---|---|
| Input | $10 | $60 |
| Cached input | $1 | $6 |
| Output | $50 | $300 |
| Bedrock US geographic routing, output | $55 | $330 |
The Bedrock row matters if you buy through AWS. The GPT-6 Astra model card on Amazon Bedrock lists Ultrafast at $66 input and $330 output for In-Region and US geographic cross-Region inference, a 10% premium over OpenAI's own rates, and $60 and $300 on Global cross-Region inference.
For scale, VentureBeat reports Ultrafast costs "three times as much as Fast mode." Fast is OpenAI's middle tier, so by VentureBeat's figures it costs twice Standard. So the step from Fast to Ultrafast is the expensive one.
Here is illustrative arithmetic, not a measured workload. A team whose agents produce one million Astra output tokens a working day pays $50 for them at Standard and $300 at Ultrafast. Over 22 working days that is $5,500 a month of premium for one million daily output tokens, before any input or cache charges, which carry the same 6x.
Where Can Ultrafast Run, and Where Can't It?
Ultrafast runs in the US or on global processing, and nowhere else. OpenAI's Ultrafast mode guide says it "supports US data residency and global processing only. It does not support EU or other non-US regional processing endpoints." The Codex and ChatGPT side says the same thing in admin language: per OpenAI's speed settings documentation, "Ultrafast isn't available to workspaces that require inference residency outside the United States."
Bedrock doesn't get you around it. On bedrock-mantle Ultrafast is available only in us-east-1, and on bedrock-runtime it is reachable through US geographic routing or Global cross-Region inference, per the same model card. AWS's own description of Global routing is that it "routes requests anywhere in the world. Use it when you have no data residency needs." An EU team can technically reach Ultrafast that way, and in doing so gives up the residency commitment it bought Standard Astra for.
This is the same pattern as the speed tier below it. The Fast mode guide says Fast "is not available with EU data residency" for GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol or GPT-6 Luna. If you run a European workspace, the premium tiers are a US product for now. Our review of OpenAI's Dots beta found the same gap on the agent side.
API capacity is also deliberately thin. The Ultrafast guide sets default limits at 500,000 tokens per minute for usage tiers 1 to 3, 1 million for tier 4 and 5 million for tier 5, and calls the current availability "low rate limits." The tier is built for a few latency-critical paths. Anything that tries to put a whole agent fleet on it will hit those limits.
How Does Ultrafast Burn Codex and Enterprise Budgets?
Inside Codex and ChatGPT Work, Ultrafast drains included usage at 8x the Standard rate and purchased credits at 6x. OpenAI's pricing documentation lists both multipliers, and the speed page spells out the split: "Ultrafast uses included subscription limits at 8x the Standard rate. Purchased credits and Enterprise pay-as-you-go usage are billed at 6x the Standard rate." Fast mode, for comparison, is 2.5x and 2x.
The two numbers mean an engineer on a fixed allowance runs it down eight times faster in Ultrafast, and anything past the allowance then bills at 6x through credits.
Enterprise eligibility is narrow. Ultrafast is limited to Pro $500, eligible Enterprise and Edu plans, and the speed page says "Legacy Enterprise plans that rely on rate limits instead of usage-based billing aren't supported." If your contract still runs on fixed rate limits, you can't turn it on; if it runs on credits or USD usage, you can.
The default works in your favour. "For Enterprise workspaces, Ultrafast is off by default. Workspace owners can enable access for selected users or the workspace," the speed documentation says. OpenAI's Enterprise usage limits page adds that "existing per-user spend limits apply to eligible Ultrafast usage" and warns that higher usage rates "can consume a user's budget faster." It doesn't describe a separate Ultrafast cap, a workspace-level Ultrafast budget or an Ultrafast-specific alert. As documented today, the per-user spend limit is the only brake.
That mirrors how usage-billed AI products have been sold all year. Microsoft billed Copilot Autopilot outside the $30 seat, and our agentic pricing guide argued then that consumption without a cap is the contract term to refuse. Ultrafast is that same consumption pricing with a 6x multiplier on top.
Which Workflows Justify Paying for Latency?
Pay for Ultrafast only where a person is waiting on generated text and their time costs more than the tokens. OpenAI's own pitch, per VentureBeat, is "interactive coding, customer support, financial analysis, incident response and other human-in-the-loop systems." That list is a reasonable starting point. Check each item against the caveat OpenAI printed itself.
The speed documentation says Ultrafast "generates tokens up to 8x faster than GPT-6 Astra in Standard mode in Codex," then adds: "This comparison measures token generation speed, not billing rates or overall task completion time." An agent turn is more than generation. It reads files, calls tools, runs tests, waits on CI and waits on a human to approve. Ultrafast speeds up the generation and leaves everything else as slow as before. A coding task that spends most of its wall-clock in a test suite gets little back for a 6x bill.
The strongest case for it looks like this:
- An on-call engineer in an active incident, reading Astra's analysis of logs while customers are down. Minutes there cost far more than tokens.
- A live support or sales conversation where the customer is on the line and Astra's answer is the bottleneck. The voice agent cost math already shows how quickly per-minute economics shift when a model is in the loop.
- Pair-programming sessions where a senior engineer reads the output as it streams and can't do anything else until it finishes.
The weakest cases are batch jobs, overnight agents, CI-triggered code review, background research and anything a person checks later. For those, the cheaper move is often a different model. GPT-6.1 Sol came in at about a fifth of Astra's price with a measurable score gap, and that trade is usually better than paying Astra 6x to finish sooner.
The strongest argument on the other side: engineer time is expensive, and an hour of a senior engineer's waiting can pay for a lot of tokens. That holds for the person watching the stream. It doesn't hold for the agent that nobody is watching, which is where most token volume goes.
What Changed for Pro 200 Holders?
The same DevDay that added Ultrafast cut the $200 Pro plan in half. Engadget reports that Pro 200's Codex and Work usage dropped from 20x the Plus allowance to 10x, and its weekly GPT-6 Pro message cap from 200 to 100. Existing subscribers keep their current limits temporarily and then receive a one-time credit. Pro 500, at $500 a month, is the plan that includes Ultrafast.
If you reimburse individual Pro subscriptions for engineers, the price of the same allowance just doubled. You can let those engineers upgrade to Pro 500, move them into the Enterprise workspace where per-user spend limits apply, or accept the halved limits. Pick before the temporary grace period runs out and the cost turns up on the next expense report.
What to Do Before Someone Turns It On
This Week:
- Open the Codex and ChatGPT workspace settings and confirm Ultrafast is still off. Then write down who holds workspace-owner rights, because those are the people who can enable it.
- Check your contract type. If you're on a legacy rate-limited Enterprise plan, Ultrafast isn't available to you, and the question waits until renewal.
- Check residency. If your workspace requires inference outside the US, Ultrafast is unavailable by design, and any team using it through a personal Pro 500 plan or Bedrock Global routing is sending that data somewhere your policy doesn't allow.
This Month:
- Set per-user spend limits for anyone you enable, sized for the 8x burn on included usage. Our prompt caching cost analysis is a useful companion: cached input also carries the 6x, so a missed cache costs proportionally more here.
- Pick two named workflows, an incident channel and one live support queue for example, and run them on Ultrafast for two weeks with a matched Standard control. Measure task completion time, not tokens per second.
- In the API, block
service_tier: "ultrafast"at your gateway for every project except the ones that passed that test. A single config change in a shared agent template can switch a whole fleet to 6x.
Before Renewal:
- Ask OpenAI for a workspace-level Ultrafast budget and an alert threshold. As of October 6, 2026, the documentation only describes per-user limits.
- Ask for an EU processing date for Ultrafast and Fast, in writing, if you have European teams.
The Bottom Line
Ultrafast is a fair product with a narrow use. OpenAI priced it linearly: up to six times the speed in the API for six times the money. It also documented its own limits honestly, from the residency gap to the caveat that faster tokens aren't faster tasks. The risk for a buyer is diffusion: a setting enabled for the one team with a latency problem gets copied into shared agent templates, and the only per-tier control OpenAI documents is a per-user spend limit, so the first signal is often the invoice.
Enable it per named user, cap it per user, and require a measured completion-time gain before it goes into any shared template.
Continue Reading
- GPT-6.1 Sol's $5.47 Task Scores 11 Points Below Astra's $23.80
- OpenAI Pulled GPT-6.1 Astra Over a Flaw GPT-6 Astra Still Has
- OpenAI's Dots Beta Skips Data Residency and Your OTel Collector
- @ChatGPT in Slack and Teams Gives Unlicensed Staff a Shared Account
- Microsoft's New Copilot Bills Autopilot Outside the $30 Seat
- Agentic AI Pricing: Don't Buy Consumption Without a Cap
