GPT-6.1 Sol is the cheapest credible way to run most agentic coding work OpenAI sells today, but it does not match GPT-6 Astra everywhere. On OpenAI's own preliminary numbers, it ties Astra on coding and costs $5.47 per science task against Astra's $23.80. It also solves 11 fewer points' worth of those science tasks. If you route traffic to Astra or Claude Opus 5.5 because of quality, the move is to re-run your own bake-off at matched scores before the next commit. Moving everything on the launch chart alone is the mistake.
OpenAI released GPT-6.1 Sol at DevDay on September 29, one week after GPT-6 Sol. It kept the $2 input / $10 output per million tokens price, which is one-fifth of Astra's $10/$50. Every benchmark behind the headline is OpenAI's, and OpenAI calls them preliminary.
Where Does GPT-6.1 Sol Actually Match Astra?
It matches Astra on coding and document work, and falls behind on computer use, workflows and science. The spread across benchmarks is the story, and the averaged "near-Astra" claim hides it.
These are the scores OpenAI published, as tabulated by Handy AI. The Opus 5.5 GDP.pdf figure (with fallbacks) is OpenAI's, via Vellum, and its Terminal-Bench Science score is from the public leaderboard:
| Benchmark (effort) | GPT-6.1 Sol | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1 (high) | 75.2% | 74.1% | — |
| GDP.pdf (high) | 32.0% | 32.2% | 28.8% |
| OSWorld 2.0 (max) | 71.4% | 73.5% | — |
| AutomationBench 1.0.6 (max) | 36.1% | 41.4% | 42.5% |
| Terminal-Bench Science 0.1 (max) | 57.0% | 68.1% | 63.3% |
On coding, Sol ties Astra and beats GPT-6 Sol by 6.4 points. On computer use it sits within 2.1 points of Astra at about one-seventh the per-task cost. OpenAI's AutomationBench claim, 2.2 points ahead of Opus 5.5 at about a third of the cost, is measured at medium effort. The max-effort row in the table shows Opus ahead. Read which effort level each number comes from before you quote it. We made the same point about GPT-6 Sol's launch evals a week ago.
Terminal-Bench Science is the widest gap. It is a 70-task benchmark of real computational workflows across five scientific domains, run by Stanford, the Laude Institute and domain experts. Sol scores 57.0% there. Astra scores 68.1% and Opus 5.5 scores 63.3%.
Is $5.47 per Task Actually Cheaper?
Yes, even after you adjust for the lower score. But that adjustment is the part that decides your workload. At max effort, OpenAI reports an average Terminal-Bench Science cost of $5.47 per task for Sol, against $23.21 for Opus 5.5 and $23.80 for Astra.
Cost per task is the price of one attempt. Cost per solved task is the price of one attempt divided by the success rate. That second number is the one your budget actually pays. Using the published scores (the costs are OpenAI's; the Opus score is the leaderboard's Claude Code run, so treat its row as an approximation):
- GPT-6.1 Sol: $5.47 ÷ 0.570 ≈ $9.60 per solved task
- GPT-6 Astra: $23.80 ÷ 0.681 ≈ $34.95 per solved task
- Claude Opus 5.5: $23.21 ÷ 0.633 ≈ $36.67 per solved task
On that math Sol still costs roughly a quarter as much per success. The strongest case for moving is exactly this: even with the gap included, the bill falls by about 70%.
The counter-case is the tasks Sol doesn't solve. On a 70-task set, an 11-point gap is about eight more failures. The savings only hold if a failure is cheap: it gets retried, escalated to a stronger model, or caught by a test. If a failure means a wrong number in a regulatory filing or a broken production migration, the roughly $18 you saved on the attempt is small next to the cost of the failure. We found the same shape in our Sonnet 5.5 vs Opus 5.5 analysis: the cheaper token can be the more expensive task.
Independent numbers are not in yet. When checked on launch day, the Terminal-Bench Science leaderboard listed Astra first at 68.1% and Opus 5.5 second at 63.3%, with no GPT-6.1 Sol entry.
What Changed Besides the Price?
Sol's safety numbers moved a long way, and the API surface has constraints that will break some integrations. Both matter more to a platform team than one more benchmark point.
On OpenAI's agentic safety evals, Sol tried to get around blocks in 23.5% of cases, down from 64.4% for GPT-6 Sol. Astra's rate is 17.4%. Unwanted outcomes, such as unauthorized transactions, fell from 17.4% to 4.3%. Astra's rate is 2.9%. So Sol is much better than its predecessor and still behind Astra. That matters for any agent with write access. Handy AI also notes a coding-deception rate of 1.50%, up from the predecessor's 1.30%, and advises against using the model for security work. Context matters here too: OpenAI reportedly scrapped the GPT-6.1 Astra release over safety concerns, including higher deception and unauthorized task execution. That is why there is a 6.1 Sol and no 6.1 Astra.
Factual accuracy improved most at low effort. The share of responses with a factual error fell from 11.4% to 7.7%, and Sol stays within 1.9 points of Astra at every reasoning setting.
On integration, reasoning is always on, and tool calling requires the Responses API. Chat Completions is supported only without tools. If your agent framework still calls Chat Completions with tool definitions, Sol is not a drop-in model swap. It needs a client migration. One pricing detail helps anyone reusing long system prompts: cached input costs $0.10 per million, 95% below the uncached rate.
What From DevDay 2026 Can You Use Today?
GPT-6.1 Sol and dots ship now, Private Intelligence is a preview, and Ultrafast on Sol is still days away. Plan against the first group only.
Generally available today:
- GPT-6.1 Sol in the API, plus ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users. It is not yet in regular ChatGPT chat.
- Dots, always-on agents that run on GPT-6 Astra with their own cloud computers and browsers. They are rolling out to Pro 200 and Business Premium, and Enterprise, Edu and Healthcare workspaces get them after an admin turns them on. Note the model: dots use Astra, not the cheaper Sol.
- Ultrafast on Astra, at 6x the standard API price, up to six times faster in the API and eight times faster in Codex. It is limited to Pro 500 and Enterprise for now.
Preview: OpenAI Private Intelligence, for keeping your data private, alongside Private Safety Processing, which runs misuse checks without breaking zero data retention. Both are previews and not a contract term yet. Do not change your data-handling posture until the terms are published.
Promised "within days": an Ultrafast version of Sol for Codex.
Seat changes worth a finance review: a new $500/month Pro tier at 25x the Plus allowance. The $200 tier's Codex/Work allowance falls from 20x to 10x Plus, and grandfathering ends October 29, 2026.
What Should Platform Teams Do Now?
Run a matched-score bake-off on your own tasks, and move only the traffic whose failures are cheap.
This Week:
- Pull the 50 most recent tasks your router sent to Astra or Opus 5.5. Replay them on GPT-6.1 Sol at the effort level you actually run in production, not max.
- Score cost per solved task, not cost per call, using the same division shown above. If your harness shaped the ranking, re-check the harness before blaming the model.
- Grep your agent code for Chat Completions calls that pass tools. Each one is a migration ticket before Sol can take that traffic.
This Month:
- Move coding and document-extraction traffic, where Sol ties Astra, behind a canary with automatic escalation to Astra on test failure.
- Keep science, long computer-use and write-access workflows on the stronger model until an independent Terminal-Bench Science entry appears.
- Get your ChatGPT seat owner to model the Pro 200 allowance cut before the October 29 grandfathering ends.
Before Renewal:
- Write "per solved task at matched score" into your model-vendor scorecard. A flat token price can still hide a doubled cost per task, and a vendor's launch chart will not show you which way yours moves.
The Bottom Line
GPT-6.1 Sol moves the price-performance line, but for now the only chart showing that is OpenAI's. Cloud buyers learned the same lesson with every new instance generation: a better list price on the launch slide says nothing about how your workload will perform. The teams that saved money re-benchmarked their own jobs before migrating.
A fifth of the price is a real number. So are eight more failed tasks out of seventy. Measure both before you move anything.
Continue Reading
- GPT-6 Sol Halves the Token Price but Benchmarks It at Top Effort
- Claude Sonnet 5.5 Costs More Than Opus 5.5 at the Same Score
- Grok 4.7 Kept Its $2/$6 Price and Can Still Double Cost per Task
- Claude Opus 5.5 Rejects the Forced Tool Calls Opus 5 Accepted
- Astra for Law Scored 54% on a Set Vals Never Publishes
- Your Bake-Off Ranked Harnesses, Not Models. Re-Run It.
