Solar Pro 4
by Upstage
Agentic enterprise LLM tuned for multi-step document work at a tenth of frontier-model token cost
Solar Pro 4 is Upstage's closed-weight commercial flagship LLM, built for agentic workloads that actually have to finish: document understanding, information extraction, long-context reasoning and multi-turn tool use. It targets enterprise teams in regulated, document-heavy industries who need instruction-following that survives many turns and tool calls that stay in valid schema, at roughly a tenth of frontier-model token cost.
Solar Pro 4 is Upstage's closed-weight flagship large language model, announced on 11 August 2026 and positioned for agentic production work rather than benchmark leaderboards. Upstage serves it through an OpenAI-compatible endpoint at api.upstage.ai/v1 under the model id solar-pro4, with reasoning enabled by default and an adjustable effort dial that trades latency against answer quality per call. The vendor documents a 512K-token context window with 128K maximum output and native English, Korean and Japanese support for both input and output, though independent measurement by Artificial Analysis recorded roughly 384K effective context rather than the advertised figure. Artificial Analysis scored the model 42 on its Intelligence Index four days after launch, up from 14 for the previous generation, placing it above NVIDIA Nemotron 3 Ultra at 38 and Google Gemini 3.5 Flash-Lite at 37. The gains cluster on the axes that break agents in production: long-context reasoning moved from 31 to 71, terminal tasks from 12 to 57, and multi-turn tool use from 9 to 23. List pricing is $0.30 per million input tokens, $0.06 cached, and $1.20 per million output tokens, with a 90% launch promotion running through 10 September 2026. Distribution spans Upstage Console, OpenRouter — where it passed 370 billion tokens in its first week — the AWS, Azure and Snowflake marketplaces, and dedicated or on-premises deployment via direct sales. Upstage, a Seoul-based unicorn founded in 2020 holding SOC 2, HIPAA, ISO 27001 and ISO 27701 certifications, sells into insurance, healthcare, manufacturing and financial services.
The platform or applied-AI lead running document-heavy agent workloads in a regulated industry, who is watching per-token spend as volume scales past pilot.
Frontier-adjacent agent reliability — 42 on Artificial Analysis's Intelligence Index — at $0.30/$1.20 per million tokens, roughly a tenth of frontier-model pricing.
At a Glance
- Category
- AI Models & APIs
- Pricing
- Usage-based, Subscription, Contact for pricing
- Target Market
- CTOs, CIOs, Enterprise Developers, Data Scientists, Heads of AI
- Deployment
- API-based, Cloud-first, Self-hosted
- Founded
- 2020
- Headquarters
- Seoul, South Korea
- Team Size
- 101-250
- Customers
- 100+ enterprise implementations (vendor-stated); customers include Samsung
Key Features
- ✓Adjustable reasoning effort dial
Reasoning is on by default with a tunable effort setting, so teams trade latency against answer quality per call.
- ✓512K-token context window
Upstage documents 512K input and 128K output tokens; independent testing measured roughly 384K effective, so verify against your longest documents.
- ✓Agentic reliability tuning
Trained so instruction-following survives multiple turns and tool calls stay in valid schema across long multi-step workflows.
- ✓Prompt caching at an 80% discount
Cached input bills at $0.06 per million tokens versus $0.30, materially cutting cost for repeated system prompts and documents.
- ✓Native English, Korean and Japanese
Handles East Asian business and regulatory documents natively, unusual among models at this price and capability tier.
- ✓Multi-channel distribution
Available through an OpenAI-compatible API, OpenRouter, the AWS, Azure and Snowflake marketplaces, plus on-premises deployment via direct sales.
Capabilities
Use Cases
- •Insurance claims document processing
Extract structured fields from claim packets and policy documents, then route decisions through multi-step agent workflows at low token cost.
- •Long-context contract and filing review
Feed hundreds of pages into a single call and rely on long-context reasoning that scored 71 on Artificial Analysis's AA-LCR benchmark.
- •Korean and Japanese enterprise document AI
Process East Asian regulatory filings and business paperwork natively rather than through a translation layer that loses formatting and nuance.
- •Agent prototyping under a hard cost ceiling
Run high-volume agent experiments where per-token spend dominates, then graduate only the workloads that genuinely need a frontier model.
- •On-premises deployment for restricted data
Run the model inside your own environment via direct sales when residency rules forbid sending documents to a public API.
Ideal For
Best For
- ✓Enterprise agent workloads where per-token cost, not peak capability, is the binding constraint
- ✓Document understanding and information extraction in insurance, healthcare, manufacturing and financial services
- ✓Long-context tasks that span multiple documents inside a single call
- ✓Teams needing native Korean or Japanese handling alongside English without a translation layer
- ✓Organisations requiring on-premises or dedicated model deployment under SOC 2 and ISO 27001 controls
Not Ideal For
- ✗Teams that need open weights to self-host or fine-tune freely — Solar Pro 4 is closed, and Upstage points those buyers at its open Solar Open 2 model instead
- ✗Latency-sensitive interactive products: Artificial Analysis measured roughly 8.6 minutes per Intelligence Index task with reasoning on by default
- ✗Genuinely frontier problems — at 42 on the Intelligence Index it sits below the top frontier models, and Upstage's own pitch is to save those models for those problems
- ✗Anyone budgeting on the launch promotional rate, which expires 10 September 2026 and reverts to ten times the price
Integrations
Deployment
Market & Ratings
100+ enterprise implementations (vendor-stated); customers include Samsung
Market Analysis
Pros
- ✓Independent third-party benchmarking from Artificial Analysis landed within days of launch, rather than vendor-only numbers
- ✓Aggressive published list pricing at $0.30/$1.20 per million tokens with an 80% cached-input discount
- ✓Large measured gains on exactly the axes that break agents in production — long-context reasoning, terminal tasks and multi-turn tool use
- ✓SOC 2, HIPAA, ISO 27001 and ISO 27701 certified vendor with on-premises deployment available through direct sales
Cons
- ✗Independent measurement found roughly 384K effective context against the 512-524K Upstage advertises
- ✗Reasoning-by-default is slow: about 8.6 minutes per Intelligence Index task in Artificial Analysis testing
- ✗Part of the measured knowledge improvement reflects the model refusing to answer rather than knowing more
- ✗Weights are closed, so anyone needing self-hosting has to drop to the separate Solar Open 2 model
- ✗Several launch benchmarks are in-house and asterisked by the vendor; independent verification is still thin outside Artificial Analysis
- ✗Almost no practitioner discussion on Hacker News for the Pro line, so production war stories are hard to find
Pricing
Pay-as-you-go API
From $0.30/1M input tokens
- ✓$0.30 per 1M input tokens
- ✓$0.06 per 1M cached input tokens
- ✓$1.20 per 1M output tokens
- ✓90% launch discount through 10 Sep 2026
Explore
From $100/mo
- ✓+10% bonus credits monthly, +30% yearly
- ✓Community support
Build
From $500/mo
- ✓+15% bonus credits monthly, +35% yearly
- ✓Email support with 3-day response
Scale
From $5,000/mo
- ✓+20% bonus credits monthly, +40% yearly
- ✓Priority support with 1-day response
Enterprise
Contact for pricing
- ✓Custom credit terms
- ✓Dedicated support
- ✓Dedicated or on-premises deployment
List pricing is fully published and metered per token with no seat licences: $0.30 per million input tokens, $0.06 cached, $1.20 output. A 90% launch promotion cuts that to $0.03/$0.12 through 10 September 2026 across both Upstage Console and OpenRouter, so budget on list rather than promo. Commitment tiers from $100 to $5,000+ per month buy bonus credits and faster support; dedicated or on-premises deployment is Enterprise-only and quoted by sales. All published prices exclude 10% Korean VAT.
Security & Compliance
Connect
Sources
This page was written from 7 sources, 5 on domains other than upstage.ai.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
OpenRouter
One OpenAI-compatible API that routes every request across 500+ models and 80+ inference providers
River AI
Token-metered LoRA fine-tuning and reinforcement learning on open-weight models you keep
NVIDIA Nemotron 3.5 Lightning
Open 30B mixture-of-experts model tuned for the high-volume execution layer of long-running agents
Scaled Cognition
APT, a large action model trained to take actions instead of predicting text