Mistral Large
by Mistral AI
Mistral's flagship model line, now a 675B Apache-2.0 mixture-of-experts at $0.50 per million input tokens
Mistral Large is the flagship general-purpose model family from French AI lab Mistral AI. The current generation, Mistral Large 3, is an Apache 2.0 mixture-of-experts model with a 256k context window and multimodal input, priced at $0.50 per million input tokens - roughly an order of magnitude below comparable US frontier APIs - and available as downloadable weights for fully on-premise deployment.
Mistral Large is the flagship general-purpose model line from French AI lab Mistral AI, and the name now refers to Mistral Large 3 (API name mistral-large-3-25-12), released on 2 December 2025. Large 3 is a granular mixture-of-experts model with 675 billion total parameters and 41 billion active per token, a 256,000-token context window, and native multimodal input: it accepts text and images and returns text. Critically for enterprise buyers, Mistral shipped it under Apache 2.0 - a genuinely permissive open-source licence with no monthly-active-user ceiling and no separate commercial agreement. That is a sharp break from its predecessor. Mistral Large 2 (123 billion dense parameters, 128k context, released 24 July 2024, 84.0% MMLU, over 80 programming languages, parallel and sequential function calling) was released under the Mistral Research License, which required a paid commercial agreement for any self-deployed business use; Artificial Analysis now marks Large 2 deprecated in favour of Large 3. On Mistral's own API, Large 3 lists at $0.50 per million input tokens and $1.50 per million output tokens - roughly 80% and 90% below GPT-5.4 on input and output respectively by CloudZero's comparison, and 83%/90% below Claude Sonnet 4.6. It supports function calling, structured outputs, document Q&A, batching, conversations and built-in tools. Distribution runs through la Plateforme, Le Chat, and the major cloud marketplaces including Azure AI, Amazon Bedrock, Google Cloud and IBM watsonx.ai; because the weights are open it can also run entirely on-premise, which is the reason European banks and regulated buyers shortlist it over US-hosted-only alternatives.
European and regulated enterprises that need frontier-adjacent model quality under EU jurisdiction, with the option to pull the weights in-house rather than depend on a US-hosted API.
Frontier-class capability at roughly a tenth of GPT-5.4's token price, under an Apache 2.0 licence that permits unrestricted commercial self-hosting.
At a Glance
- Category
- AI Models & APIs
- Pricing
- Usage-based, Subscription, Freemium, Contact for pricing
- Target Market
- CIOs, CTOs, Enterprise Developers, Data Protection Officers, Heads of AI
- Deployment
- API-based, Open-source, Self-hosted, Multi-cloud, Hybrid
- Founded
- 2023
- Headquarters
- Paris, France
- Team Size
- 201-500
Key Features
- ✓Granular mixture-of-experts architecture
675B total parameters with only 41B active per token, so serving cost tracks the small number while capability tracks the large one.
- ✓256,000-token context window
Handles book-length contracts, full repositories or large retrieval sets in one call without an external chunking pipeline.
- ✓Apache 2.0 open weights
No MAU ceiling and no commercial agreement required, which is materially more permissive than the community licences on competing open-weight models.
- ✓Native multimodal input
Accepts images alongside text, so document, diagram and screenshot workflows do not need a separate vision model.
- ✓Function calling and structured outputs
Supports parallel and sequential tool calls with schema-constrained JSON, which is the prerequisite for reliable agent orchestration.
- ✓Multi-cloud and on-premise distribution
Available on la Plateforme, Azure AI, Amazon Bedrock, Google Cloud and IBM watsonx.ai, or self-deployed from the open weights.
Capabilities
Use Cases
- •Sovereign enterprise assistant
European banks and insurers deploy the open weights in-country so no customer data crosses a jurisdictional boundary during inference.
- •High-volume document extraction
Process contracts, claims and invoices at $0.50 per million input tokens where a frontier US API would cost roughly ten times more.
- •Multilingual customer operations
Handle French, German, Spanish, Italian and Portuguese support queues on one model rather than one vendor per language.
- •Tool-calling agent backends
Drive internal automations using parallel function calls and structured outputs, with the 256k window holding full tool schemas and history.
- •Code assistance across large repositories
The Large 2 lineage trained on 80+ programming languages, and the long context lets a whole service be reasoned about in one prompt.
Ideal For
Best For
- ✓EU-headquartered enterprises with data-residency or sovereignty requirements that rule out US-only model hosting
- ✓High-volume text workloads where a 10x token-price difference compounds into a material line item
- ✓Regulated industries - banking, insurance, healthcare - that need on-premise deployment of a capable general model
- ✓Multilingual European document processing across French, German, Spanish, Italian and Portuguese
- ✓Agentic and tool-calling backends needing parallel function calls and guaranteed structured output
Not Ideal For
- ✗Teams that want the single highest benchmark score regardless of cost - Large 3 is priced as a value flagship, and dedicated reasoning models still lead on hard maths and competitive coding
- ✗Anyone planning to self-host casually: 675 billion total parameters means a serious multi-GPU cluster despite the permissive Apache 2.0 licence, so 'open weights' does not translate into 'runs on a workstation'
- ✗Buyers who bought into Mistral Large 2 expecting a long support horizon - it was superseded and marked deprecated inside about seventeen months, and its research licence made self-deployment a paid negotiation
- ✗Le Chat subscribers expecting API access; the consumer and developer products are entirely separate billing streams
Integrations
Deployment
Market Analysis
Pros
- ✓Apache 2.0 licensing on a 675B flagship removes the MAU ceilings and commercial-agreement negotiations that encumber most open-weight models
- ✓Published list pricing at $0.50/$1.50 per million tokens undercuts comparable US frontier APIs by roughly 80-90%
- ✓256k context plus native image input covers long-document and screenshot workflows without bolting on a second model
- ✓Available simultaneously on Azure AI, Amazon Bedrock, Google Cloud, IBM watsonx.ai and as raw weights, so there is a genuine exit from any single host
Cons
- ✗Short support horizon on this line: Mistral Large 2 shipped July 2024 and Artificial Analysis already marks it deprecated in favour of Large 3, forcing a migration inside about seventeen months
- ✗Large 2's Mistral Research License required a paid commercial agreement for self-deployment, so buyers who invested in that generation did not get the open-source terms Large 3 now advertises
- ✗Apache 2.0 weights are only useful to teams that can host 675B total parameters - for everyone else the licence is theoretical and the API price is the real cost
- ✗Le Chat and the API are separate billing streams, and a Pro subscription grants no API access at all, which CloudZero flags as a recurring source of buyer confusion
- ✗Artificial Analysis placed Large 2 at #13 of 39 in its open-weight non-reasoning cohort - respectable rather than leading, and dedicated reasoning models still outperform this line on hard problems
Pricing
API - Mistral Large 3
From $0.50 per 1M input tokens
- ✓$0.50/M input, $1.50/M output
- ✓256k context window
- ✓Function calling and structured outputs
- ✓Batching and built-in tools
Open weights (Apache 2.0)
$0
- ✓Download and self-host with no per-token fee
- ✓No MAU ceiling
- ✓Unrestricted commercial use
- ✓Requires substantial multi-GPU capacity
Le Chat Pro
From $14.99/mo
- ✓Extended thinking
- ✓Deep research
- ✓Consumer assistant only - does not include API access
Le Chat Enterprise
Contact for pricing
- ✓SAML SSO
- ✓White label
- ✓Cloud marketplace and on-premise deployment
Mistral publishes list API pricing, which is unusual at this tier: Large 3 is metered at $0.50 per million input tokens and $1.50 per million output, which CloudZero measures at roughly 80% cheaper on input and 90% cheaper on output than GPT-5.4, and 83%/90% cheaper than Claude Sonnet 4.6. The Apache 2.0 weights make the per-token cost avoidable entirely if you have the GPUs, but 675B total parameters means that is a cluster decision rather than a workstation one. Two billing traps matter: Le Chat subscriptions (Free, Pro $14.99/mo, Team $24.99/user/mo) are a completely separate stream from API credits and buy no API access, and Le Chat Enterprise with SAML SSO and white-labelling is quoted rather than listed.
Security & Compliance
Connect
Sources
This page was written from 5 sources, 4 on domains other than mistral.ai.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Scaled Cognition
APT, a large action model trained to take actions instead of predicting text
Engram
A learned memory layer that makes AI actually know your organization — at up to 100x fewer tokens.
TwelveLabs
Video intelligence API that makes every hour of enterprise footage searchable, analyzable and agent-ready.
Mistral OCR 4
Structure-aware document AI that returns bounding boxes, typed blocks, and per-word confidence scores.
Mentioned In
AWS vs GCP vs Azure ML: The Real Costs Nobody Tells You
Enterprise AI analysis: AWS vs GCP vs Azure. Strategic insights, ROI considerations, and implementation guidance for technical and business leaders evaluatin...
March 15, 2026Enterprise AIAI Enterprise Adoption Hits Inflection Point in Q1 2026
Enterprise AI analysis: AI Enterprise Adoption Hits Inflection Point in Q1 2026. Strategic insights, ROI considerations, and implementation guidance for tech...
March 26, 2026Enterprise AIVoice AI's Rent vs. Own Moment: Mistral Bets on Open Weights
Mistral's open-weight Voxtral TTS challenges the $22B voice AI market with a bet that enterprises will choose data sovereignty over convenience.
March 30, 2026