Sakana AI Fugu Max
by Sakana AI
Multi-agent orchestration behind one API — frontier-grade results at $2 per million input tokens
Fugu Max is Sakana AI's cost-optimised orchestration model: rather than one monolithic network, it routes each sub-task to the leanest capable model in a swappable pool of open-weight and specialist models, then synthesises the result. It is exposed through OpenAI- and Anthropic-compatible APIs for enterprises that want frontier-level output without frontier-level token spend or single-vendor dependency.
Fugu Max, released on 11 September 2026 alongside Fugu Ultra v2.0, is the cost-optimised tier of Sakana AI's Fugu orchestration architecture. Fugu is not a single model but a model interface: an orchestration layer decomposes an incoming request into sub-tasks, routes each to the leanest capable model in a swappable pool of open-weight and specialised models — NVIDIA Nemotron models among them — verifies the outputs and synthesises a final answer. Sakana frames the design as supply-chain resilience: the pool is vendor-agnostic, enterprises can exclude specific providers from routing for compliance reasons, and Sakana notes Fugu Ultra v2 reaches its scores without Fable 5, Fable 5.1 or GPT-6-Astra in its agent pool. Fugu Max is priced at $2 per million input tokens, $6 per million output and $0.25 per million cached input, which Sakana says undercuts Sonnet 5, GPT-5.6 Terra and Kimi K3 by 40-60% on output; Fugu Ultra v2 runs $5/$30/$0.50 with higher rates above a 272K-token context. Both are reachable through OpenAI-compatible Responses and Chat Completions APIs and an Anthropic-compatible Messages API using the identifiers fugu-max and fugu-ultra, so existing integrations switch with a single parameter change. Sakana reports Fugu Max taking top scores on six benchmarks including Terminal Bench 2.1, GPQAD, AutomationBench and SWEFish. Sakana AI was founded in Tokyo in July 2023 by David Ha (CEO), Ren Ito (Chairman) and Llion Jones (CTO, a co-author of 'Attention Is All You Need'), has raised over $368 million at a roughly $2.7 billion valuation, and deploys sovereign AI systems at Japanese institutions including MUFG and Daiwa Securities. Fugu is hosted API only and is not available in the EU or EEA.
The head of AI platform at a non-EU enterprise running high-volume agentic or analytical workloads who needs frontier-class output but cannot justify frontier per-token pricing or accept single-vendor model dependency.
Frontier-comparable benchmark results at $2/$6 per million tokens — vendor-reported 40-60% below Sonnet 5, GPT-5.6 Terra and Kimi K3 on output — reachable through a one-line API switch.
At a Glance
- Category
- AI Models & APIs
- Pricing
- Usage-based, Subscription
- Target Market
- CTOs, Heads of AI, Enterprise Developers, ML Engineers, Data Scientists
- Deployment
- API-based, Cloud-only
- Founded
- 2023
- Headquarters
- Tokyo, Japan
Key Features
- ✓Dynamic task routing
Decomposes a request and sends each sub-task to the leanest model capable of solving it, cutting token spend materially.
- ✓Swappable heterogeneous agent pool
Mixes open-weight and specialised models including NVIDIA Nemotron, so no single provider is load-bearing for availability.
- ✓Provider exclusion for compliance
Enterprises can remove specific model providers from the routing pool to satisfy procurement or regulatory constraints.
- ✓Dual API compatibility
Speaks OpenAI Responses and Chat Completions plus the Anthropic Messages API, so existing clients migrate with one parameter change.
- ✓Two cost/capability tiers
fugu-max optimises cost at $2/$6 per million tokens while fugu-ultra maximises capability at $5/$30 on the same architecture.
- ✓Prompt caching and long context
Cached input drops to $0.25 per million tokens, with a 272K-token standard context band before higher long-context rates apply.
Capabilities
Use Cases
- •Cost reduction on an existing agent stack
Point an OpenAI-compatible agent framework at fugu-max to cut output token cost 40-60% without rewriting client code.
- •Document and chart analysis at volume
Sakana reports best-in-class scores on GDP.pdf and Chartography, where Fugu Ultra v2 scores 48.3 against competitors' 27.3-29.5.
- •Terminal and automation agents
Fugu Max leads Terminal Bench 2.1 and AutomationBench, the workloads closest to enterprise back-office task execution.
- •Vendor-risk hedging for AI platforms
Routing across a heterogeneous pool insulates production from one provider's outage, price change or export restriction.
- •Software engineering assistance
Fugu Ultra v2 scores 74.3 on DeepSWE and leads SWEFish, targeting code-heavy agent loops at lower cost than frontier models.
Ideal For
Best For
- ✓High-volume agentic workloads where per-token cost, not peak capability, is the binding constraint
- ✓Teams wanting provider redundancy built into the model layer so an outage or export restriction does not halt production
- ✓Compliance-driven routing where specific model providers must be excluded from the pool
- ✓Drop-in cost reduction for existing OpenAI- or Anthropic-API integrations via a single parameter change
- ✓Document, chart and terminal-style automation tasks, where Sakana reports its strongest benchmark results
Not Ideal For
- ✗Any organisation serving EU or EEA users — Fugu is explicitly unavailable in those markets pending GDPR alignment, which rules it out for most European enterprises outright
- ✗Buyers who need to know and pin exactly which model handled a request: the worker pool and routing logic are proprietary and deliberately hidden, a point researcher Elie Bakouch raised directly against Sakana's sovereignty framing
- ✗Workloads demanding single-model determinism or precise cost prediction — Fugu Ultra exposes extra orchestration token fields that add cost beyond visible prompt and output tokens, and parameters such as temperature are accepted and ignored
- ✗Air-gapped or self-hosted requirements: Fugu is hosted API only, with no on-premise or open-weights option
- ✗Highest-stakes design and reasoning work where hands-on comparisons found Fugu finished faster but with minor logic errors against Claude Opus 4.8's slower, stronger output
Integrations
Deployment
Market Analysis
Pros
- ✓Genuinely aggressive economics for the claimed capability tier — $2/$6 per million tokens, vendor-reported 40-60% below Sonnet 5, GPT-5.6 Terra and Kimi K3 on output, with cached input at $0.25
- ✓Architectural hedge against vendor risk: a swappable heterogeneous pool, per-customer provider exclusion, and demonstrated frontier-level results without the leading proprietary models in the pool
- ✓Zero-friction adoption for existing stacks — OpenAI Responses/Chat Completions and Anthropic Messages compatibility means a single-line parameter change
- ✓Serious institutional backing and real enterprise deployment: $368M+ raised at ~$2.7B, investors including Google, NVIDIA and Salesforce Ventures, and production work at MUFG and Daiwa Securities
Cons
- ✗Not available in the EU or EEA pending GDPR alignment, which disqualifies it outright for most European enterprises regardless of price or benchmark
- ✗The routing logic and worker-model pool are proprietary and hidden by design; research engineer Elie Bakouch's criticism — 'you don't even control which ones are used or how much' — cuts directly against the sovereignty pitch
- ✗Developers characterise Fugu as a highly advanced router or wrapper rather than a fundamental capability advance, and the benchmarks are vendor-reported with noted competitor exclusions
- ✗Cost is harder to predict than the price list implies: Fugu Ultra exposes extra orchestration token fields that add to the bill beyond visible prompt and output tokens
- ✗Hands-on comparison found Fugu completed complex tasks faster but with minor logic errors while Claude Opus 4.8 was slower and produced the better final design, and some API parameters such as temperature are accepted then ignored
- ✗Hosted API only, no self-hosting or open weights, with a stated training cutoff of 28 August 2026
Pricing
Fugu Max (pay-as-you-go)
From $2 per 1M input tokens
- ✓$6 per 1M output tokens
- ✓$0.25 per 1M cached input tokens
- ✓$0.007 per web_search or web_fetch call
- ✓OpenAI- and Anthropic-compatible APIs
Fugu Ultra v2 (pay-as-you-go)
From $5 per 1M input tokens
- ✓$30 per 1M output tokens
- ✓$0.50 per 1M cached input tokens
- ✓$10/$45/$1 per 1M above 272K context
- ✓Highest-capability tier
Subscription
From $20/mo
- ✓Tiers reported from $20 to $200 per month
- ✓For everyday rather than high-volume use
Token pricing is published: Fugu Max is $2 input / $6 output / $0.25 cached per million tokens and Fugu Ultra v2 is $5 / $30 / $0.50, rising to $10 / $45 / $1 above a 272K-token context, with web_search and web_fetch tool calls billed at $0.007 each. Subscriptions are reported from $20 to $200 a month for lighter use. The important caveat for budgeting is that Fugu Ultra exposes additional orchestration token fields that add to total cost beyond the visible prompt and output counts, so an effective rate cannot be derived from the headline price alone.
Security & Compliance
Sources
This page was written from 5 sources, 3 on domains other than sakana.ai.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Arcee Trinity
US-built open-weight model family, from on-device Trinity Nano to the 400B-parameter Trinity Large, that you can run on your own infrastructure
Abacus.AI Smaug
Open-weight Smaug Agentic, Flash and Mini models fine-tuned for long-running enterprise AI agents
Deep Cogito Cogito v2.1
MIT-licensed 671B hybrid-reasoning open model with short reasoning chains, plus custom post-training on enterprise data
GPT-6 Astra
OpenAI's frontier model for autonomous computer use, gated cyber capability and long-horizon coding