Kimi K3
by Moonshot AI
2.8T-parameter open-weight MoE model with a 1M-token context, priced at $3/$15 per million tokens
Kimi K3 is Moonshot AI's open-weight frontier language model: a 2.8 trillion parameter mixture-of-experts system with 104B active parameters, a 1,048,576-token context window and native image and video understanding. It is for engineering and AI platform teams that want near-frontier coding and agent performance with the option to download the weights and run them on their own GPUs.
Kimi K3 is a large language model from Beijing-based Moonshot AI, launched on July 16, 2026, with weights published on Hugging Face under moonshotai/Kimi-K3 about eleven days later. The model is a mixture-of-experts design with 2.8 trillion total parameters and 104 billion activated per token, using 896 experts (16 routed plus 2 shared) and a hybrid attention stack of 69 KDA layers and 24 gated MLA layers. It supports a 1,048,576-token context and handles text, images and video in one model. Moonshot trained it with quantization-aware training from supervised fine-tuning onward, so the released checkpoint uses MXFP4 weights and MXFP8 activations, roughly 1.56 TB across 96 shards. The GitHub model card reports 93.5 on GPQA Diamond, 88.3 on Terminal-Bench 2.1, 91.2 on BrowseComp and 67.5 on DeepSWE; Artificial Analysis placed it third among open models on its Intelligence Index (43.6) as of September 2026. It can be served with vLLM, SGLang or TokenSpeed, and is available through Moonshot's OpenAI- and Anthropic-compatible API at $3.00 input, $0.30 cached input and $15.00 output per million tokens, as well as the Kimi app, Kimi Code and OpenRouter. The weights ship under a custom Kimi K3 License: internal use is unrestricted, but a company with more than $20 million in trailing 12-month revenue that offers K3 to third parties as a model service must sign a separate agreement with Moonshot first. Buyers should also weigh slow inference (about 37 tokens per second in third-party testing), capacity limits that led Moonshot to pause new subscriptions days after launch, and ongoing US congressional scrutiny of American companies using Kimi models.
A head of AI platform or ML infrastructure who needs a near-frontier coding and agent model they can self-host for internal workloads, and who has the multi-node GPU capacity and legal review to do it.
Near-frontier coding and long-context agent performance with downloadable weights, or API access at $3/$15 per million tokens, below Claude Opus 5.5's $4/$20 list price.
At a Glance
- Category
- AI Models & APIs
- Pricing
- Usage-based
- Target Market
- CTOs, Heads of AI, ML Platform Engineers, Enterprise Developers
- Deployment
- API-based, Self-hosted
- Founded
- 2023
- Headquarters
- Beijing, China
- Team Size
- 201-500
Key Features
- ✓2.8T-parameter mixture-of-experts
896 experts with 16 routed and 2 shared per token keep active compute at 104B parameters, which holds serving cost well below a dense model of the same size.
- ✓1,048,576-token context
Lets an agent hold an entire codebase or a long document set in one prompt, and Moonshot's API charges the same rate across the full window with no long-context surcharge.
- ✓Native multimodal input
Text, image and video understanding sit in one model, with reported scores of 94.3 on MathVision and 90.0 on Video-MME, so visual tasks do not need a second model.
- ✓MXFP4 quantization-aware training
Weights are trained for MXFP4 with MXFP8 activations from fine-tuning onward, so the released low-precision checkpoint is the one that was evaluated rather than a lossy post-hoc compression.
- ✓Open weights with vLLM and SGLang support
The checkpoint is downloadable from Hugging Face and runs on vLLM, SGLang or TokenSpeed, so teams can self-host for internal use without paying per token.
- ✓Two-tier prompt caching
Cache writes cost $3.00 (5-minute TTL) or $6.00 (1-hour TTL) per million tokens, and cached reads cost $0.30, a 90% discount that matters for agents re-sending large contexts.
Capabilities
Use Cases
- •Repository-scale coding agents
Load a full monorepo into the 1M-token window so an agent can plan cross-file refactors and run terminal tasks, where K3 reports 88.3 on Terminal-Bench 2.1.
- •Self-hosted internal assistant
Run the open weights on an in-house GPU cluster so employee prompts and documents never leave the company network, which the licence permits for internal use without restriction.
- •Deep research and browsing agents
Use K3 as the reasoning model behind a web research agent, where it reports 91.2 on BrowseComp, to compile sourced briefings from many pages.
- •Cost-tiered model routing
Route long-context and coding traffic to K3 at $3/$15 per million tokens while reserving pricier closed models for tasks where they measurably win, lowering the blended inference bill.
Ideal For
Best For
- ✓Internal coding agents that need a 1M-token context to load large repositories without chunking
- ✓Teams that must keep inference on their own GPU clusters for data-control reasons and can host a 1.56 TB checkpoint
- ✓Front-end and web development generation, where K3 was the first open-weight model to top the Arena WebDev leaderboard
- ✓Benchmarking and red-teaming an open frontier model alongside closed US models before choosing a production default
Not Ideal For
- ✗Inference providers or SaaS vendors above $20M in annual revenue who want to resell K3 as a model service, because the licence requires a separate commercial agreement with Moonshot first
- ✗Latency-sensitive interactive apps, since third-party tests measured roughly 37 tokens per second, the slowest frontier model in that comparison
- ✗Regulated US organisations whose policies or customers restrict Chinese-developed models, given congressional scrutiny of US firms using Kimi models
- ✗Small teams without multi-node GPU clusters who want to self-host at full context
Integrations
Deployment
Market Analysis
Pros
- ✓Open weights at near-frontier quality: third among open models on the Artificial Analysis Intelligence Index (43.6) and first open-weight model on Arena WebDev
- ✓API pricing about 25% below Claude Opus 5.5's list price, with 90%-discounted cached input
- ✓Unrestricted internal use under the licence, so self-hosting for employees needs no contract
- ✓Strong practitioner reception on Hacker News, where the launch thread drew over 2,100 points
Cons
- ✗Slow: third-party testing measured about 37 tokens per second, and Hacker News users reported long waits even for simple code reviews
- ✗Capacity-constrained at launch: Moonshot paused new subscriptions in July 2026, and users complained that the $20/month tier's layered quotas ran out within a day
- ✗The Kimi K3 License is a commercial instrument, not a permissive open-source grant; reselling it as a service above $20M revenue requires a deal with Moonshot
- ✗Jurisdiction and reputational risk: US lawmakers have scrutinised American firms using Kimi models, and Anthropic has accused Moonshot of distilling Claude outputs through fraudulent accounts
- ✗A UK AISI / CAISI preliminary assessment found it 'significantly below' leading US models on offensive cyber tasks (0 of 41 code-execution samples versus about 20)
Pricing
Kimi API (kimi-k3)
$3.00 input / $15.00 output per 1M tokens
- ✓$0.30 per 1M cached input tokens
- ✓Cache writes $3.00 (5-min TTL) or $6.00 (1-hour TTL) per 1M tokens
- ✓1,048,576-token context
- ✓OpenAI- and Anthropic-compatible endpoints
Open weights (self-hosted)
$0
- ✓Download from Hugging Face under the Kimi K3 License
- ✓Unrestricted internal use
- ✓Separate agreement required to offer it as a model service above $20M annual revenue
- ✓Attribution required above 100M MAU or $20M monthly revenue
API list price is $3.00 per million input tokens, $0.30 cached and $15.00 output, flat across the 1M context (Moonshot pricing page, checked October 2026); reasoning is always on and thinking tokens bill as output, so real bills run higher than prompt length suggests. Self-hosting the weights costs nothing in licence fees for internal use but needs a multi-node GPU cluster, and commercial resale above $20M revenue requires a negotiated contract.
Security & Compliance
Connect
Sources
This page was written from 9 sources, 9 on domains other than kimi.com.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Xiaomi MiMo-V2.6
MIT-licensed, 1M-context multimodal MoE models that took the top open-weights score on Artificial Analysis at well under $1 per million tokens
TypeSafe Jev
A 'System One' decision model that returns typed choices, scores and calibrated probabilities instead of text, for agent routing and classification
PrismML Bonsai
Open-weight 1-bit and ternary LLMs that run 27B-class AI on laptops and phones
Sakana AI Fugu Max
Multi-agent orchestration behind one API — frontier-grade results at $2 per million input tokens