A

Anthropic Claude Opus 4.6

by Anthropic

AI Models & APIsAI Agents & OrchestrationDeveloper ToolsEnterprise Platform

The February 2026 Claude flagship that introduced adaptive thinking and effort control

Usage-based · Contact for pricing·Added Mar 14, 2026·Updated Aug 2, 2026
Share:
THE DAILY BRIEF
Anthropic Claude Opus 4.6

by Anthropic

AI Models & APIsAI Agents & OrchestrationDeveloper ToolsEnterprise Platform

The February 2026 Claude flagship that introduced adaptive thinking and effort control

Usage-based · Contact for pricing

Claude Opus 4.6 is Anthropic's February 2026 frontier model, still active on the API as a pinned snapshot at $5 per million input and $25 per million output tokens. It was the first Claude release with adaptive thinking and a four-level effort control, pairing a 1M-token context with 128K output for long-running agentic coding, research and professional knowledge work.

At a Glance

Category
AI Models & APIs
Pricing
Usage-based, Contact for pricing
Target Market
CTOs, CIOs, Enterprise Developers, Data Scientists, AI Research Engineers
Deployment
API-based, Cloud-only, Multi-cloud
Founded
2021
Headquarters
San Francisco, United States
Team Size
500+

Key Features

  • Adaptive thinking
  • Four-level effort control
  • 1M-token context window
  • 128K output, 300K in batch
  • Prompt caching
  • Data residency via inference_geo
  • Multi-cloud availability

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • Autonomous agentic coding
  • Financial analysis and research
  • Deep web research
  • Frozen evaluation baselines
  • Long-document professional knowledge work

Ideal For

Best For

  • Long-running autonomous coding agents that operate across a whole workflow rather than a single file — METR estimated a 50%-success time horizon of around 14.5 hours
  • Whole-repository and whole-corpus reasoning where the 1M-token context removes the retrieval layer entirely
  • Reasoning-heavy analysis where the effort parameter can be pushed to max with a large thinking budget, as Anthropic did for its ARC-AGI-2 runs
  • Regression-sensitive production pipelines that need a pinned model ID whose behaviour will not shift under them until at least February 2027
  • Codebases still passing temperature, top_p or top_k, which return a 400 error on Claude Opus 4.7 and later

Not Ideal For

  • Latency-sensitive interactive products — BenchLM measures about 40 output tokens per second against a 108 field median, and unlike Opus 5 and Opus 4.8 it cannot use Fast mode, which runs at standard speed and standard rates if requested
  • Anyone starting fresh today: Opus 4.7, Opus 4.8 and Opus 5 all sit above it at the identical $5/$25 price, so choosing 4.6 for new work buys older capability for the same money
  • High-volume or cost-sensitive workloads — at $25 per million output tokens it is five times Haiku 4.5, and thinking tokens bill as output
  • Work that depends on recent events: its reliable knowledge cutoff is May 2025, and Simon Willison's system-prompt analysis shows Anthropic had to hard-code current-affairs corrections into the prompt to compensate

Market Analysis

Enterprise-gradeFrontier modelAgentic coding

Pros

  • Frontier-tier scores that still hold up months later — 80.8% SWE-bench Verified, 91.3% GPQA and 82% MMLU-Pro on BenchLM's independent leaderboard, ranking #20 of 214 models
  • 1M-token context with 128K output (300K in batch) removes the retrieval layer for whole-repository and whole-corpus work
  • Adaptive thinking plus a four-level effort dial gives per-request control over the accuracy-versus-cost trade rather than a single global setting
  • Anthropic has committed not to retire it before 5 February 2027 and gives at least 60 days' notice, which is what a pinned production dependency needs

Cons

  • Slow: BenchLM measures about 40 output tokens per second against a 108 field median, and Fast mode — available on Opus 5 and 4.8 — is explicitly unsupported here
  • Superseded three times over. Opus 4.7, Opus 4.8 and Opus 5 all sit above it at the exact same $5/$25 price, so there is no economic reason to pick it for new work
  • Reliable knowledge cutoff of May 2025 is the weakest of the 4.6-and-later generation; Anthropic's own system prompt had to hard-code corrections about current events to compensate
  • The explicit extended-thinking parameter is marked deprecated on this model, so code written against it is already on a migration path
  • Reliability wobbles are on the public record — Anthropic's status page logged an 'Elevated errors on Claude Opus 4.6' incident, and practitioners posted a BridgeBench run showing hallucination accuracy dropping from 83% to 68%
  • Price pressure from below is real: HN threads point to MiniMax M2.5 claiming comparable SWE-bench performance at roughly 20x lower cost

Pricing

Claude API (pay-as-you-go)

From $5/M input tokens

  • $5 per million input tokens
  • $25 per million output tokens
  • 5-minute cache writes $6.25/M, cache reads $0.50/M
  • Full 1M context at standard rates

Batch API

From $2.50/M input tokens

  • 50% discount on both input and output
  • $2.50/M input, $12.50/M output
  • Supports the 300K extended-output beta header
  • Not compatible with Fast mode

Cloud marketplaces

Contact for pricing

  • Amazon Bedrock and Google Cloud billed by the provider
  • Claude Platform on AWS and Microsoft Foundry bill in Claude Consumption Units at $0.01 per CCU
  • US-only inference at a 1.1x multiplier via inference_geo
  • Private offers and negotiated volume discounts

Purely token-metered at $5 input and $25 output per million, with no seat licence — the same rate as Opus 4.5, Opus 4.7, Opus 4.8 and Opus 5, which is the key commercial fact: the whole Opus line is priced identically, so 4.6 offers no discount for being a generation behind. The launch-era long-context premium is gone; Anthropic now bills the full 1M window at standard rates, and prompt caching drops reads to $0.50 per million. Batch halves both sides to $2.50/$12.50. Fast mode is not available on 4.6 — requests with `speed: "fast"` simply run at standard speed and standard rates rather than erroring. US-only inference via `inference_geo` adds a 1.1x multiplier across every token category, and regional endpoints on Bedrock or Google Cloud add 10% over global. Nothing is gated behind an Enterprise SKU; custom rate limits and volume discounts are negotiated with sales.

Security & Compliance

soc2
hipaa
iso27001
data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, twice weekly.

beri.net

Subscribe at beri.net/subscribe for twice-weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Claude Opus 4.6 is Anthropic's February 2026 frontier model, still active on the API as a pinned snapshot at $5 per million input and $25 per million output tokens. It was the first Claude release with adaptive thinking and a four-level effort control, pairing a 1M-token context with 128K output for long-running agentic coding, research and professional knowledge work.

Claude Opus 4.6 is Anthropic's February 2026 frontier release, announced on 5 February 2026 and still listed as an Active model on the Claude API under the ID `claude-opus-4-6`, with a tentative retirement date no sooner than 5 February 2027. It was the first Claude model to ship adaptive thinking — where the model reads contextual cues about how much extended reasoning a request warrants — alongside an explicit `effort` parameter offering low, medium, high and max levels, and it pairs a 1M-token context window with up to 128,000 output tokens, rising to 300,000 on the Batch API behind the `output-300k-2026-03-24` beta header. Anthropic reported state-of-the-art results at launch on Terminal-Bench 2.0, Humanity's Last Exam, GDPval-AA and BrowseComp, citing 81.42% on SWE-bench Verified with prompt modification, 53.0% on Humanity's Last Exam with tools, and roughly a 144-Elo lead over GPT-5.2 on GDPval-AA. Independent aggregation is more measured: BenchLM records 80.8% SWE-bench Verified, 91.3% GPQA and 82% MMLU-Pro, ranks it #20 of 214 models overall and #8 of 55 on knowledge, but clocks throughput at about 40 output tokens per second against a field median of 108. Pricing is $5 per million input tokens and $25 per million output — unchanged from Opus 4.5 and identical to today's Opus 5 — and the full 1M context now bills at standard rates rather than the launch-era long-context premium. It runs on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic has since shipped Opus 4.7, Opus 4.8 and Opus 5 above it, so 4.6 now serves as a price-stable pinned snapshot rather than the recommended default.

Ideal Buyer

Engineering and research teams that pinned an agent or evaluation harness to a specific early-2026 frontier snapshot and need it to keep behaving identically — plus anyone who still relies on the temperature, top_p and top_k sampling controls that Claude 4.7 and later reject outright.

Key Benefit

Frontier-grade agentic coding and long-horizon reasoning across a 1M-token context at $5/$25 per million tokens, on a snapshot Anthropic has committed not to retire before 5 February 2027.

At a Glance

Category
AI Models & APIs
Pricing
Usage-based, Contact for pricing
Target Market
CTOs, CIOs, Enterprise Developers, Data Scientists, AI Research Engineers
Deployment
API-based, Cloud-only, Multi-cloud
Founded
2021
Headquarters
San Francisco, United States
Team Size
500+

Key Features

  • Adaptive thinking

    The model infers from context how much extended reasoning a request needs, instead of requiring a manually set thinking budget on every call

  • Four-level effort control

    Low, medium, high and max settings let you trade latency and token spend against depth explicitly per request

  • 1M-token context window

    Holds roughly 750,000 words in a single request, now billed at standard per-token rates across the whole window

  • 128K output, 300K in batch

    Emits up to 128,000 tokens synchronously and 300,000 via the Batch API extended-output beta header

  • Prompt caching

    Cache reads bill at $0.50 per million tokens, a tenth of base input price, which matters most for long system prompts

  • Data residency via inference_geo

    US-only inference is selectable at a 1.1x pricing multiplier, a control unavailable on models older than the 4.6 generation

  • Multi-cloud availability

    Ships on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS with regional endpoints

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • Autonomous agentic coding

    Runs multi-hour terminal and repository tasks end to end, the workload Anthropic tuned it for and measured on Terminal-Bench 2.0

  • Financial analysis and research

    Anthropic published a dedicated finance-focused Opus 4.6 programme; long context suits filings, transcripts and model documentation in one pass

  • Deep web research

    Anthropic reported category-leading BrowseComp results, so multi-hop browsing agents that must chase citations across sites are a natural fit

  • Frozen evaluation baselines

    Research teams comparing models over time need a pinned snapshot whose scores stay reproducible for the length of a study

  • Long-document professional knowledge work

    Contract review, technical due diligence and specification analysis where the 1M context avoids chunking and retrieval error

Ideal For

Best For

  • Long-running autonomous coding agents that operate across a whole workflow rather than a single file — METR estimated a 50%-success time horizon of around 14.5 hours
  • Whole-repository and whole-corpus reasoning where the 1M-token context removes the retrieval layer entirely
  • Reasoning-heavy analysis where the effort parameter can be pushed to max with a large thinking budget, as Anthropic did for its ARC-AGI-2 runs
  • Regression-sensitive production pipelines that need a pinned model ID whose behaviour will not shift under them until at least February 2027
  • Codebases still passing temperature, top_p or top_k, which return a 400 error on Claude Opus 4.7 and later

Not Ideal For

  • Latency-sensitive interactive products — BenchLM measures about 40 output tokens per second against a 108 field median, and unlike Opus 5 and Opus 4.8 it cannot use Fast mode, which runs at standard speed and standard rates if requested
  • Anyone starting fresh today: Opus 4.7, Opus 4.8 and Opus 5 all sit above it at the identical $5/$25 price, so choosing 4.6 for new work buys older capability for the same money
  • High-volume or cost-sensitive workloads — at $25 per million output tokens it is five times Haiku 4.5, and thinking tokens bill as output
  • Work that depends on recent events: its reliable knowledge cutoff is May 2025, and Simon Willison's system-prompt analysis shows Anthropic had to hard-code current-affairs corrections into the prompt to compensate

Integrations

SDK Available
SDK:PythonC#

Deployment

On-Premise

Market Analysis

Enterprise-gradeFrontier modelAgentic coding

Pros

  • Frontier-tier scores that still hold up months later — 80.8% SWE-bench Verified, 91.3% GPQA and 82% MMLU-Pro on BenchLM's independent leaderboard, ranking #20 of 214 models
  • 1M-token context with 128K output (300K in batch) removes the retrieval layer for whole-repository and whole-corpus work
  • Adaptive thinking plus a four-level effort dial gives per-request control over the accuracy-versus-cost trade rather than a single global setting
  • Anthropic has committed not to retire it before 5 February 2027 and gives at least 60 days' notice, which is what a pinned production dependency needs

Cons

  • Slow: BenchLM measures about 40 output tokens per second against a 108 field median, and Fast mode — available on Opus 5 and 4.8 — is explicitly unsupported here
  • Superseded three times over. Opus 4.7, Opus 4.8 and Opus 5 all sit above it at the exact same $5/$25 price, so there is no economic reason to pick it for new work
  • Reliable knowledge cutoff of May 2025 is the weakest of the 4.6-and-later generation; Anthropic's own system prompt had to hard-code corrections about current events to compensate
  • The explicit extended-thinking parameter is marked deprecated on this model, so code written against it is already on a migration path
  • Reliability wobbles are on the public record — Anthropic's status page logged an 'Elevated errors on Claude Opus 4.6' incident, and practitioners posted a BridgeBench run showing hallucination accuracy dropping from 83% to 68%
  • Price pressure from below is real: HN threads point to MiniMax M2.5 claiming comparable SWE-bench performance at roughly 20x lower cost

Pricing

Free Trial Available

Claude API (pay-as-you-go)

From $5/M input tokens

  • $5 per million input tokens
  • $25 per million output tokens
  • 5-minute cache writes $6.25/M, cache reads $0.50/M
  • Full 1M context at standard rates

Batch API

From $2.50/M input tokens

  • 50% discount on both input and output
  • $2.50/M input, $12.50/M output
  • Supports the 300K extended-output beta header
  • Not compatible with Fast mode

Cloud marketplaces

Contact for pricing

  • Amazon Bedrock and Google Cloud billed by the provider
  • Claude Platform on AWS and Microsoft Foundry bill in Claude Consumption Units at $0.01 per CCU
  • US-only inference at a 1.1x multiplier via inference_geo
  • Private offers and negotiated volume discounts

Purely token-metered at $5 input and $25 output per million, with no seat licence — the same rate as Opus 4.5, Opus 4.7, Opus 4.8 and Opus 5, which is the key commercial fact: the whole Opus line is priced identically, so 4.6 offers no discount for being a generation behind. The launch-era long-context premium is gone; Anthropic now bills the full 1M window at standard rates, and prompt caching drops reads to $0.50 per million. Batch halves both sides to $2.50/$12.50. Fast mode is not available on 4.6 — requests with `speed: "fast"` simply run at standard speed and standard rates rather than erroring. US-only inference via `inference_geo` adds a 1.1x multiplier across every token category, and regional endpoints on Bedrock or Google Cloud add 10% over global. Nothing is gated behind an Enterprise SKU; custom rate limits and volume discounts are negotiated with sales.

Security & Compliance

soc2
hipaa
iso27001
data residency

Connect

Sources

This page was written from 9 sources, 8 on domains other than anthropic.com.

  1. 1.anthropic.comclaude opus 4 6vendor
  2. 2.platform.claude.compricing
  3. 3.platform.claude.commodel deprecations
  4. 4.platform.claude.comoverview
  5. 5.benchlm.aiclaude opus 4 6
  6. 6.openrouter.aiclaude opus 4.6
  7. 7.hn.algolia.comhn.algolia.com
  8. 8.simonwillison.netopus system prompt
  9. 9.en.wikipedia.orgAnthropic
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →