O

OpenAI o3

by OpenAI

AI Models & APIsDeveloper Tools

OpenAI's 2025 reasoning model — still serving, but scheduled for shutdown on 11 December 2026

Usage-based·Added Mar 14, 2026·Updated Aug 3, 2026
Share:
THE DAILY BRIEF
OpenAI o3

by OpenAI

AI Models & APIsDeveloper Tools

OpenAI's 2025 reasoning model — still serving, but scheduled for shutdown on 11 December 2026

Usage-based

o3 is OpenAI's April 2025 reasoning model for maths, science, coding and visual reasoning, with a 200,000-token context window. It remains callable today but is formally deprecated: OpenAI announced its retirement on 11 June 2026 with a shutdown date of 11 December 2026 and gpt-5.6-sol as the named replacement.

At a Glance

Category
AI Models & APIs
Pricing
Usage-based
Target Market
CTOs, Enterprise Developers, Data Scientists, Platform Engineers
Deployment
API-based, Cloud-only
Founded
2015
Headquarters
San Francisco, United States
Team Size
500+

Key Features

  • Reasoning across maths, science and coding
  • 200,000-token context window
  • Prompt caching at 75 percent discount
  • Structured outputs and function calling
  • Image input and visual reasoning
  • Evals and stored completions

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • Migration benchmarking to gpt-5.6-sol
  • Scientific and mathematical reasoning
  • Visual reasoning over technical diagrams
  • Legacy production traffic during wind-down
  • Cost-controlled batch reasoning

Ideal For

Best For

  • Auditing existing production workloads that still pin o3-2025-04-16 or o3-pro-2025-06-10 before the December 2026 shutdown
  • Short-lived or fixed-term projects that will complete well before 11 December 2026 and already have o3 prompts tuned
  • Cost-sensitive reasoning work needing high output throughput, where 124.7 tokens per second at $8 per million output is competitive
  • Benchmarking a gpt-5.6-sol migration against a known o3 baseline on the same evaluation set
  • Workloads exploiting the 75 percent prompt-cache discount, at $0.50 per million cached input against $2.00 standard

Not Ideal For

  • Any new build — the model shuts down on 11 December 2026, so a greenfield integration starts with a migration debt already on the books
  • Long-context work: the 200,000-token window is a fifth of GPT-5.4's and half of GPT-5 mini's, at a higher output price than the latter
  • Tasks needing current knowledge, given a 1 June 2024 cutoff that is now over two years old
  • Value-sensitive buyers comparing on measured capability — Artificial Analysis rates it at intelligence index 30, 97th of 185 and below average, while charging $8 per million output
  • Agentic desktop or computer-use work, which o3 does not support and which is native to the GPT-5.4 line

Market Analysis

Enterprise-gradeLegacyDeprecated

Pros

  • Still fully callable with documented pricing and features, and OpenAI has published a clear replacement in gpt-5.6-sol
  • Artificial Analysis measures 124.7 tokens per second output, fast for its price band
  • Prompt caching at $0.50 per million is a 75 percent discount on standard input
  • Broad feature support including structured outputs, function calling, image input, evals and stored completions
  • Time-to-first-token of 5.07 seconds is far better than the frontier GPT-5.4 line at high reasoning effort

Cons

  • Deprecated: announced 11 June 2026, shutdown 11 December 2026 — every workload on it needs migrating within months
  • The family is being dismantled piecewise, with o3-deep-research already shut down on 23 July 2026 and o3-mini shutting down 23 October 2026
  • Artificial Analysis rates it below average at intelligence index 30, 97th of 185 models, while charging $8 per million output
  • The 200,000-token context window is a fifth of GPT-5.4's and half of GPT-5 mini's
  • Knowledge cutoff of 1 June 2024 is over two years stale
  • No computer use, agentic desktop control or workflow automation — capabilities that moved to the GPT-5.4 line

Pricing

Standard API (until 11 December 2026)

From $2.00 per 1M input tokens

  • $2.00 per 1M input tokens
  • $0.50 per 1M cached input tokens
  • $8.00 per 1M output tokens
  • 200,000-token context window, 100,000 max output
  • Shutdown scheduled for 11 December 2026

Per-token metering at $2.00 input, $0.50 cached and $8.00 output per million, with no seats or commitments. The real cost consideration is not the rate but the clock: OpenAI announced retirement on 11 June 2026 with shutdown on 11 December 2026, so any budget built on o3 pricing must also carry the engineering cost of migrating to gpt-5.6-sol and re-validating output quality before that date.

Security & Compliance

soc2
gdpr
hipaa
iso27001
sso
data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, twice weekly.

beri.net

Subscribe at beri.net/subscribe for twice-weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

o3 is OpenAI's April 2025 reasoning model for maths, science, coding and visual reasoning, with a 200,000-token context window. It remains callable today but is formally deprecated: OpenAI announced its retirement on 11 June 2026 with a shutdown date of 11 December 2026 and gpt-5.6-sol as the named replacement.

o3 is OpenAI's reasoning model released on 16 April 2025, documented as a well-rounded and powerful model across domains that set a new standard for maths, science, coding and visual reasoning. It accepts text and image input and returns text, with a 200,000-token context window, up to 100,000 output tokens and a knowledge cutoff of 1 June 2024. Supported features include streaming, structured outputs, function calling, file search, file uploads, image input, prompt caching, evals and stored completions. List pricing is $2.00 per million input tokens, $0.50 per million cached input and $8.00 per million output. The single most important fact for anyone evaluating it today is that o3 is on OpenAI's published deprecation schedule: the retirement of o3-2025-04-16 and o3-pro-2025-06-10 was announced on 11 June 2026 with a shutdown date of 11 December 2026, and gpt-5.6-sol is the named replacement, with reasoning.mode set to pro for the o3-pro path. The wider family is already being dismantled — o3-deep-research shut down on 23 July 2026 and o3-mini-2025-01-31 shuts down on 23 October 2026, both also replaced by gpt-5.6-sol. Independent measurement by Artificial Analysis places o3 at an intelligence index of 30, ranking 97th of 185 models and characterised as below average in intelligence but reasonably priced against models in the same band, with output at 124.7 tokens per second and time-to-first-token of 5.07 seconds. In practice o3 is now a legacy endpoint: it is genuinely fast for its price and its prompt-cache discount of 75 percent is generous, but every workload running on it needs a migration plan dated well before December 2026, and its 200,000-token window is a fifth of what GPT-5.4 offers at a comparable output price.

Ideal Buyer

Teams already running production traffic on o3 who need the migration facts — shutdown date, named replacement and what changes — rather than teams choosing a reasoning model today.

Key Benefit

A firm, published shutdown date of 11 December 2026 and a named successor in gpt-5.6-sol, so migration can be planned rather than discovered when calls start failing.

At a Glance

Category
AI Models & APIs
Pricing
Usage-based
Target Market
CTOs, Enterprise Developers, Data Scientists, Platform Engineers
Deployment
API-based, Cloud-only
Founded
2015
Headquarters
San Francisco, United States
Team Size
500+

Key Features

  • Reasoning across maths, science and coding

    Documented by OpenAI as setting a new standard on maths, science, coding and visual reasoning tasks at release.

  • 200,000-token context window

    Accepts up to 100,000 output tokens per completion, adequate for long documents though a fifth of GPT-5.4's window.

  • Prompt caching at 75 percent discount

    Cached input at $0.50 per million against $2.00 standard, a meaningful saving on repeated system context.

  • Structured outputs and function calling

    Returns schema-conformant JSON and invokes tools, so existing extraction and agent integrations need no parsing layer.

  • Image input and visual reasoning

    Analyses diagrams, screenshots and scanned documents alongside text within the same reasoning request.

  • Evals and stored completions

    Native support for OpenAI's evaluation tooling, useful for benchmarking a replacement model against a recorded o3 baseline.

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • Migration benchmarking to gpt-5.6-sol

    Run a stored eval set against both models to quantify quality and cost changes before the December 2026 cutover.

  • Scientific and mathematical reasoning

    Multi-step derivation and proof-checking work that o3 was specifically built and benchmarked for at release.

  • Visual reasoning over technical diagrams

    Interpret schematics, charts and screenshots where the reasoning depends on spatial and structural relationships.

  • Legacy production traffic during wind-down

    Keep an established, prompt-tuned pipeline serving while the replacement is validated, with a fixed end date.

  • Cost-controlled batch reasoning

    High-throughput offline analysis exploiting 124.7 tokens per second output and the 75 percent prompt-cache discount.

Ideal For

Best For

  • Auditing existing production workloads that still pin o3-2025-04-16 or o3-pro-2025-06-10 before the December 2026 shutdown
  • Short-lived or fixed-term projects that will complete well before 11 December 2026 and already have o3 prompts tuned
  • Cost-sensitive reasoning work needing high output throughput, where 124.7 tokens per second at $8 per million output is competitive
  • Benchmarking a gpt-5.6-sol migration against a known o3 baseline on the same evaluation set
  • Workloads exploiting the 75 percent prompt-cache discount, at $0.50 per million cached input against $2.00 standard

Not Ideal For

  • Any new build — the model shuts down on 11 December 2026, so a greenfield integration starts with a migration debt already on the books
  • Long-context work: the 200,000-token window is a fifth of GPT-5.4's and half of GPT-5 mini's, at a higher output price than the latter
  • Tasks needing current knowledge, given a 1 June 2024 cutoff that is now over two years old
  • Value-sensitive buyers comparing on measured capability — Artificial Analysis rates it at intelligence index 30, 97th of 185 and below average, while charging $8 per million output
  • Agentic desktop or computer-use work, which o3 does not support and which is native to the GPT-5.4 line

Integrations

SDK Available
SDK:PythonTypeScriptJavaScript

Deployment

On-Premise

Market Analysis

Enterprise-gradeLegacyDeprecated

Pros

  • Still fully callable with documented pricing and features, and OpenAI has published a clear replacement in gpt-5.6-sol
  • Artificial Analysis measures 124.7 tokens per second output, fast for its price band
  • Prompt caching at $0.50 per million is a 75 percent discount on standard input
  • Broad feature support including structured outputs, function calling, image input, evals and stored completions
  • Time-to-first-token of 5.07 seconds is far better than the frontier GPT-5.4 line at high reasoning effort

Cons

  • Deprecated: announced 11 June 2026, shutdown 11 December 2026 — every workload on it needs migrating within months
  • The family is being dismantled piecewise, with o3-deep-research already shut down on 23 July 2026 and o3-mini shutting down 23 October 2026
  • Artificial Analysis rates it below average at intelligence index 30, 97th of 185 models, while charging $8 per million output
  • The 200,000-token context window is a fifth of GPT-5.4's and half of GPT-5 mini's
  • Knowledge cutoff of 1 June 2024 is over two years stale
  • No computer use, agentic desktop control or workflow automation — capabilities that moved to the GPT-5.4 line

Pricing

Standard API (until 11 December 2026)

From $2.00 per 1M input tokens

  • $2.00 per 1M input tokens
  • $0.50 per 1M cached input tokens
  • $8.00 per 1M output tokens
  • 200,000-token context window, 100,000 max output
  • Shutdown scheduled for 11 December 2026

Per-token metering at $2.00 input, $0.50 cached and $8.00 output per million, with no seats or commitments. The real cost consideration is not the rate but the clock: OpenAI announced retirement on 11 June 2026 with shutdown on 11 December 2026, so any budget built on o3 pricing must also carry the engineering cost of migrating to gpt-5.6-sol and re-validating output quality before that date.

Security & Compliance

soc2
gdpr
hipaa
iso27001
sso
data residency

Connect

Sources

This page was written from 6 sources, 4 on domains other than developers.openai.com.

  1. 1.developers.openai.como3vendor
  2. 2.developers.openai.comdeprecationsvendor
  3. 3.artificialanalysis.aio3
  4. 4.openrouter.aio3
  5. 5.trust.openai.comtrust.openai.com
  6. 6.en.wikipedia.orgOpenAI
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe