Anthropic Claude Haiku 4.5
by Anthropic
Anthropic's fastest Claude model — near-frontier coding at $1/$5 per million tokens
Claude Haiku 4.5 is Anthropic's fastest and cheapest current Claude model, released 15 October 2025 and priced at $1 per million input tokens and $5 per million output. It gives platform and application teams near-frontier coding and computer-use quality — 73.3% on SWE-bench Verified — at roughly a third of Sonnet-class cost, so high-volume chat, sub-agent and classification workloads stop being a budget line nobody can defend.
Claude Haiku 4.5 is the small, fast tier of Anthropic's Claude family, released on 15 October 2025 and exposed through the API as `claude-haiku-4-5-20251001` (alias `claude-haiku-4-5`). It carries a 200,000-token context window, emits up to 64,000 output tokens, accepts text and images, and is the only model in Anthropic's current line-up that still uses the explicit `thinking.type: "enabled"` extended-thinking switch rather than the newer adaptive-thinking `effort` control. Anthropic positions it as near-frontier intelligence at commodity price: it scores 73.3% on SWE-bench Verified, reaches 50.7% on OSWorld computer use, and the company says it delivers roughly 90% of Claude Sonnet 4.5's agentic-coding performance while running four to five times faster. Pricing is $1 per million input tokens and $5 per million output, with cache reads at $0.10 per million and Batch API rates of $0.50/$2.50 — the arithmetic Anthropic itself uses to cost a customer-support workload at about $37 per 10,000 tickets. Independent measurement broadly agrees with the positioning: Artificial Analysis records an Intelligence Index of 24 (#29 of 72 models, against a median of 19), 89.3 output tokens per second and a 0.96-second time to first token, calling it above average in intelligence and well priced against other non-reasoning models in its band. The practical role is the sub-agent — orchestrate with a frontier model, execute with Haiku — plus free-tier features, real-time chat, classification and extraction at volume. It is available on the Claude API, claude.ai, Claude Code, Amazon Bedrock, Google Cloud Vertex AI and Microsoft Foundry, and Anthropic names GitHub, Notion, Replit, Warp and Augment among its users.
The platform engineering lead running high-volume inference — support agents, classifiers, extraction pipelines, or sub-agents inside a larger orchestration — where per-call cost and latency, not peak reasoning, decide the architecture.
Claude Sonnet 4-class coding and computer-use quality at $1/$5 per million tokens and four to five times the speed of Sonnet 4.5, which is what makes a fan-out sub-agent design affordable at all.
At a Glance
- Category
- AI Models & APIs
- Pricing
- Usage-based, Contact for pricing
- Target Market
- CTOs, Enterprise Developers, Platform Engineers, Data Scientists, Product Engineering Leads
- Deployment
- API-based, Cloud-only, Multi-cloud
- Founded
- 2021
- Headquarters
- San Francisco, United States
- Team Size
- 500+
Key Features
- ✓200K context / 64K output
A 200,000-token context window with up to 64,000 output tokens, enough for long conversations and multi-file code changes without chunking
- ✓Extended thinking
Brings controllable reasoning depth to the Haiku line via `thinking.type: enabled`, letting you trade latency for accuracy per request
- ✓Computer use
Scores 50.7% on OSWorld, the highest any Haiku model has reached, enabling supervised GUI automation at low cost
- ✓Prompt caching
Cache reads bill at $0.10 per million tokens — 10% of base input price — so repeated system prompts and documents stop dominating spend
- ✓Batch API at 50% off
Asynchronous processing at $0.50 input and $2.50 output per million tokens for work that does not need an immediate response
- ✓Multimodal input and tool use
Accepts text and images and supports tool calling, structured outputs, bash, web search and code execution tools
- ✓Multi-cloud availability
Ships on the Claude API, Amazon Bedrock, Google Cloud Vertex AI and Microsoft Foundry, with regional endpoints for data-residency requirements
Capabilities
Use Cases
- •Customer support automation
Anthropic costs a 10,000-ticket support workload at roughly $37 total using Haiku 4.5 at about 3,700 tokens per conversation
- •Multi-agent sub-task execution
A frontier model handles planning and orchestration while Haiku 4.5 executes each sub-task, cutting the cost of fan-out architectures
- •Real-time coding assistance
Powers inline completion and pair-programming surfaces where 89.3 output tokens per second keeps the interaction feeling immediate
- •Research and literature synthesis sub-agents
Runs parallel retrieval and summarisation workers under a coordinating model for literature reviews and market monitoring
- •High-volume classification and extraction
Batch-processes documents, tickets and transcripts into structured fields at $0.50 per million input tokens overnight
Ideal For
Best For
- ✓Sub-agent execution inside a multi-agent system where a frontier model plans and Haiku carries out each individual sub-task
- ✓Real-time customer service agents and chat assistants where a 0.96-second time to first token is the product
- ✓Free-tier and trial product surfaces that need usable intelligence without Sonnet- or Opus-level unit economics
- ✓High-volume classification, extraction and document triage billed at $0.50/$2.50 per million tokens through the Batch API
- ✓Coding sub-agents and pair-programming assistants that need 73.3% SWE-bench Verified quality at a third of Sonnet 4.5's cost
Not Ideal For
- ✗Long-context agents: the 200,000-token window is one-fifth of the 1M available on Claude Sonnet 4.6, Opus 4.6 and every current Claude 5 model, so large repositories and long document sets will not fit
- ✗Unattended computer-use automation — 50.7% on OSWorld means roughly half of attempts fail, and Caylent's production write-up says outright that it is not reliable enough for autonomous deployment without approval workflows
- ✗Work that depends on recent world knowledge: its reliable knowledge cutoff is February 2025, the oldest of any model Anthropic currently lists as active
- ✗Teams that want the model to scale its own reasoning depth — Haiku 4.5 has no adaptive thinking or `effort` parameter, so thinking budgets are tuned by hand
Integrations
Deployment
Market Analysis
Pros
- ✓Genuinely cheap for the capability — $1/$5 per million tokens with cache reads at $0.10 makes fan-out and free-tier architectures viable
- ✓Fast in independent measurement, not just vendor claims: Artificial Analysis records 89.3 output tokens/sec and 0.96s time to first token, both well ahead of the median for its class
- ✓73.3% on SWE-bench Verified is within five points of Sonnet 4.5 at roughly a third of the price
- ✓Runs on every major cloud — Bedrock, Vertex AI, Microsoft Foundry and Claude Platform on AWS — with regional endpoints for data residency
Cons
- ✗200,000-token context is one-fifth of the 1M window on Sonnet 4.6, Opus 4.6 and the Claude 5 line, which rules it out of long-repository and long-document agent runs
- ✗Computer use fails roughly half the time at 50.7% OSWorld; Caylent's production analysis states plainly that the 49.3% failure rate makes it unsuitable for autonomous deployment
- ✗Oldest knowledge of any active Claude model — reliable knowledge cutoff February 2025, training data cutoff July 2025
- ✗No adaptive thinking or `effort` parameter, so unlike every newer Claude model it cannot scale its own reasoning depth to the request
- ✗Its tentative retirement date is no sooner than 15 October 2026 — the nearest of any active Claude model, so pipelines pinned to it face a migration sooner than those on the 4.6 generation
- ✗Intelligence Index of 24 places it #29 of 72 on Artificial Analysis; it is priced well but is not close to frontier reasoning
- ✗25% more expensive than the Haiku 3.5 it replaced, so simple classification workloads that ran fine on 3.5 got a price rise, not a saving
Pricing
Claude API (pay-as-you-go)
From $1/M input tokens
- ✓$1 per million input tokens
- ✓$5 per million output tokens
- ✓5-minute cache writes $1.25/M, cache reads $0.10/M
- ✓1-hour cache writes $2.00/M
Batch API
From $0.50/M input tokens
- ✓50% discount on input and output
- ✓$0.50/M input, $2.50/M output
- ✓Asynchronous processing for non-latency-sensitive work
- ✓Stacks with prompt caching discounts
Cloud marketplaces
Contact for pricing
- ✓Amazon Bedrock and Google Cloud Vertex AI billed by the provider
- ✓Claude Platform on AWS and Microsoft Foundry bill in Claude Consumption Units at $0.01 per CCU
- ✓Regional and multi-region endpoints carry a 10% premium over global
- ✓Private offers and volume discounts via account teams
Metered purely per token at $1 input and $5 output per million, with no seat licence and no minimum — the cheapest active Claude model, though 25% more expensive than the retired Haiku 3.5 at $0.80/$4. Real spend is governed by caching rather than list price: cache hits bill at $0.10 per million, so a stable system prompt costs a tenth of the headline rate, and the Batch API halves both sides again to $0.50/$2.50. Server-side tools add on top — web search is $10 per 1,000 searches, web fetch is free, code execution is free when paired with web search and otherwise $0.05 per container-hour after 1,550 free hours a month. Nothing is gated behind an Enterprise tier; volume discounts and custom rate limits above the Scale tier are negotiated with sales.
Security & Compliance
Connect
Sources
This page was written from 10 sources, 8 on domains other than anthropic.com.
- 1.anthropic.com — claude haiku 4 5vendor
- 2.anthropic.com — haikuvendor
- 3.platform.claude.com — pricing
- 4.platform.claude.com — model deprecations
- 5.artificialanalysis.ai — claude 4 5 haiku
- 6.openrouter.ai — claude haiku 4.5
- 7.llm-stats.com — claude haiku 4 5 20251001
- 8.caylent.com — claude haiku 4 5 deep dive cost capabilities and the multi a
- 9.hn.algolia.com — hn.algolia.com
- 10.en.wikipedia.org — Anthropic
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Scaled Cognition
APT, a large action model trained to take actions instead of predicting text
Engram
A learned memory layer that makes AI actually know your organization — at up to 100x fewer tokens.
TwelveLabs
Video intelligence API that makes every hour of enterprise footage searchable, analyzable and agent-ready.
Mistral OCR 4
Structure-aware document AI that returns bounding boxes, typed blocks, and per-word confidence scores.