OpenAI o3
by OpenAI
OpenAI's 2025 reasoning model — still serving, but scheduled for shutdown on 11 December 2026
o3 is OpenAI's April 2025 reasoning model for maths, science, coding and visual reasoning, with a 200,000-token context window. It remains callable today but is formally deprecated: OpenAI announced its retirement on 11 June 2026 with a shutdown date of 11 December 2026 and gpt-5.6-sol as the named replacement.
o3 is OpenAI's reasoning model released on 16 April 2025, documented as a well-rounded and powerful model across domains that set a new standard for maths, science, coding and visual reasoning. It accepts text and image input and returns text, with a 200,000-token context window, up to 100,000 output tokens and a knowledge cutoff of 1 June 2024. Supported features include streaming, structured outputs, function calling, file search, file uploads, image input, prompt caching, evals and stored completions. List pricing is $2.00 per million input tokens, $0.50 per million cached input and $8.00 per million output. The single most important fact for anyone evaluating it today is that o3 is on OpenAI's published deprecation schedule: the retirement of o3-2025-04-16 and o3-pro-2025-06-10 was announced on 11 June 2026 with a shutdown date of 11 December 2026, and gpt-5.6-sol is the named replacement, with reasoning.mode set to pro for the o3-pro path. The wider family is already being dismantled — o3-deep-research shut down on 23 July 2026 and o3-mini-2025-01-31 shuts down on 23 October 2026, both also replaced by gpt-5.6-sol. Independent measurement by Artificial Analysis places o3 at an intelligence index of 30, ranking 97th of 185 models and characterised as below average in intelligence but reasonably priced against models in the same band, with output at 124.7 tokens per second and time-to-first-token of 5.07 seconds. In practice o3 is now a legacy endpoint: it is genuinely fast for its price and its prompt-cache discount of 75 percent is generous, but every workload running on it needs a migration plan dated well before December 2026, and its 200,000-token window is a fifth of what GPT-5.4 offers at a comparable output price.
Teams already running production traffic on o3 who need the migration facts — shutdown date, named replacement and what changes — rather than teams choosing a reasoning model today.
A firm, published shutdown date of 11 December 2026 and a named successor in gpt-5.6-sol, so migration can be planned rather than discovered when calls start failing.
At a Glance
- Category
- AI Models & APIs
- Pricing
- Usage-based
- Target Market
- CTOs, Enterprise Developers, Data Scientists, Platform Engineers
- Deployment
- API-based, Cloud-only
- Founded
- 2015
- Headquarters
- San Francisco, United States
- Team Size
- 500+
Key Features
- ✓Reasoning across maths, science and coding
Documented by OpenAI as setting a new standard on maths, science, coding and visual reasoning tasks at release.
- ✓200,000-token context window
Accepts up to 100,000 output tokens per completion, adequate for long documents though a fifth of GPT-5.4's window.
- ✓Prompt caching at 75 percent discount
Cached input at $0.50 per million against $2.00 standard, a meaningful saving on repeated system context.
- ✓Structured outputs and function calling
Returns schema-conformant JSON and invokes tools, so existing extraction and agent integrations need no parsing layer.
- ✓Image input and visual reasoning
Analyses diagrams, screenshots and scanned documents alongside text within the same reasoning request.
- ✓Evals and stored completions
Native support for OpenAI's evaluation tooling, useful for benchmarking a replacement model against a recorded o3 baseline.
Capabilities
Use Cases
- •Migration benchmarking to gpt-5.6-sol
Run a stored eval set against both models to quantify quality and cost changes before the December 2026 cutover.
- •Scientific and mathematical reasoning
Multi-step derivation and proof-checking work that o3 was specifically built and benchmarked for at release.
- •Visual reasoning over technical diagrams
Interpret schematics, charts and screenshots where the reasoning depends on spatial and structural relationships.
- •Legacy production traffic during wind-down
Keep an established, prompt-tuned pipeline serving while the replacement is validated, with a fixed end date.
- •Cost-controlled batch reasoning
High-throughput offline analysis exploiting 124.7 tokens per second output and the 75 percent prompt-cache discount.
Ideal For
Best For
- ✓Auditing existing production workloads that still pin o3-2025-04-16 or o3-pro-2025-06-10 before the December 2026 shutdown
- ✓Short-lived or fixed-term projects that will complete well before 11 December 2026 and already have o3 prompts tuned
- ✓Cost-sensitive reasoning work needing high output throughput, where 124.7 tokens per second at $8 per million output is competitive
- ✓Benchmarking a gpt-5.6-sol migration against a known o3 baseline on the same evaluation set
- ✓Workloads exploiting the 75 percent prompt-cache discount, at $0.50 per million cached input against $2.00 standard
Not Ideal For
- ✗Any new build — the model shuts down on 11 December 2026, so a greenfield integration starts with a migration debt already on the books
- ✗Long-context work: the 200,000-token window is a fifth of GPT-5.4's and half of GPT-5 mini's, at a higher output price than the latter
- ✗Tasks needing current knowledge, given a 1 June 2024 cutoff that is now over two years old
- ✗Value-sensitive buyers comparing on measured capability — Artificial Analysis rates it at intelligence index 30, 97th of 185 and below average, while charging $8 per million output
- ✗Agentic desktop or computer-use work, which o3 does not support and which is native to the GPT-5.4 line
Integrations
Deployment
Market Analysis
Pros
- ✓Still fully callable with documented pricing and features, and OpenAI has published a clear replacement in gpt-5.6-sol
- ✓Artificial Analysis measures 124.7 tokens per second output, fast for its price band
- ✓Prompt caching at $0.50 per million is a 75 percent discount on standard input
- ✓Broad feature support including structured outputs, function calling, image input, evals and stored completions
- ✓Time-to-first-token of 5.07 seconds is far better than the frontier GPT-5.4 line at high reasoning effort
Cons
- ✗Deprecated: announced 11 June 2026, shutdown 11 December 2026 — every workload on it needs migrating within months
- ✗The family is being dismantled piecewise, with o3-deep-research already shut down on 23 July 2026 and o3-mini shutting down 23 October 2026
- ✗Artificial Analysis rates it below average at intelligence index 30, 97th of 185 models, while charging $8 per million output
- ✗The 200,000-token context window is a fifth of GPT-5.4's and half of GPT-5 mini's
- ✗Knowledge cutoff of 1 June 2024 is over two years stale
- ✗No computer use, agentic desktop control or workflow automation — capabilities that moved to the GPT-5.4 line
Pricing
Standard API (until 11 December 2026)
From $2.00 per 1M input tokens
- ✓$2.00 per 1M input tokens
- ✓$0.50 per 1M cached input tokens
- ✓$8.00 per 1M output tokens
- ✓200,000-token context window, 100,000 max output
- ✓Shutdown scheduled for 11 December 2026
Per-token metering at $2.00 input, $0.50 cached and $8.00 output per million, with no seats or commitments. The real cost consideration is not the rate but the clock: OpenAI announced retirement on 11 June 2026 with shutdown on 11 December 2026, so any budget built on o3 pricing must also carry the engineering cost of migrating to gpt-5.6-sol and re-validating output quality before that date.
Security & Compliance
Connect
Sources
This page was written from 6 sources, 4 on domains other than developers.openai.com.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Scaled Cognition
APT, a large action model trained to take actions instead of predicting text
Engram
A learned memory layer that makes AI actually know your organization — at up to 100x fewer tokens.
TwelveLabs
Video intelligence API that makes every hour of enterprise footage searchable, analyzable and agent-ready.
Mistral OCR 4
Structure-aware document AI that returns bounding boxes, typed blocks, and per-word confidence scores.