Xiaomi MiMo-V2.6
by Xiaomi
MIT-licensed, 1M-context multimodal MoE models that took the top open-weights score on Artificial Analysis at well under $1 per million tokens
Xiaomi MiMo-V2.6 is a family of open-weight, MIT-licensed multimodal language models, led by the 1.02-trillion-parameter Pro and the 310-billion-parameter Flash, built for coding agents and professional workflows. It is for enterprises that want near-frontier agentic capability they can self-host or buy by API at well under $1 per million tokens.
MiMo-V2.6 is the September 2026 release of Xiaomi's MiMo model family, developed by a team led by Luo Fuli, formerly of DeepSeek. Released on September 21, 2026 under the MIT licence, it comprises MiMo-V2.6-Pro, a sparse mixture-of-experts model with 1.02 trillion total and 42 billion active parameters (384 routed experts, 8 active per token), and MiMo-V2.6-Flash, with 310 billion total and 15 billion active, plus a faster Pro-UltraSpeed variant and a 9B model distilled onto Qwen. Both main models accept text, image, audio and video input, offer a 1-million-token context window with up to 128,000 output tokens, and produce text. The release centres on reinforcement learning: Xiaomi says the Pro run covered roughly 750,000 agent trajectories in six days, and it has open-sourced its technical report, RL framework, more than 7,000 task environments and composable mini-harnesses alongside the weights on Hugging Face. Pro scored 46 on the Artificial Analysis Intelligence Index, tying for the top open-weights score at launch and ahead of DeepSeek V4.1 Pro and Gemini 3.8 Flash, though Claude Opus 5 still leads on DeepSWE, ProgramBench and Terminal Bench. Xiaomi's API prices Pro at $0.435 per million input tokens and $0.87 output, and Flash at $0.14 and $0.28, while a Token Plan subscription from $6 a month bundles credits for coding tools such as Claude Code, OpenCode and OpenClaw. Self-hosting Pro is a multi-node job: the model card recommends SGLang across two nodes of H100/H200-class GPUs. The line began with MiMo-7B in April 2025; the MiMo-V2-Pro flagship of March 2026 was proprietary before V2.5 and now V2.6 returned the flagship to open weights.
Heads of AI platform at cost-sensitive enterprises who want a near-frontier, permissively licensed model they can either self-host for data control or buy cheaply by API.
Top-tier open-weights intelligence (Artificial Analysis Intelligence Index 46) under the MIT licence, at API prices under $1 per million tokens.
At a Glance
- Category
- AI Models & APIs
- Pricing
- Free, Usage-based, Subscription
- Target Market
- CTOs, Heads of AI, ML Platform Engineers, Enterprise Developers
- Deployment
- Open-source, Self-hosted, API-based
- Founded
- 2010
- Headquarters
- Beijing, China
- Team Size
- 500+
Key Features
- ✓Trillion-parameter sparse MoE
Pro activates 42B of 1.02T parameters per token, balancing frontier-class capability against inference cost.
- ✓1M-token multimodal context
Pro and Flash both accept text, image, audio and video input with up to 128K output tokens for long-horizon agent tasks.
- ✓MIT-licensed open weights
Weights on Hugging Face in BF16 and FP8 permit commercial use, fine-tuning and on-premises deployment without royalties.
- ✓Open-sourced RL stack
Xiaomi released its RL framework, 7,000+ task environments and mini-harnesses, so teams can reproduce or extend agent training.
- ✓Low-cost API and Token Plan
Pay-as-you-go API from $0.14 per million input tokens, plus $6-$100 monthly credit plans aimed at coding-agent tools.
- ✓Multi-token prediction decoding
A 5-layer speculative decoder predicts 7 tokens per pass, and a Pro-UltraSpeed variant targets much faster output generation.
Capabilities
Use Cases
- •Self-hosted coding agent
Run Flash or Pro on internal GPUs behind the firewall so coding agents work on source code without sending it to a third-party API.
- •Cutting agent pipeline cost
Move high-volume agent steps from a proprietary frontier model to Flash at $0.14/$0.28 per million tokens where evaluations show quality holds.
- •Long multimodal review
Analyze long contracts, meeting recordings or video in a single 1M-token context window with one model instead of a pipeline.
- •Custom agent post-training
Use the released RL environments, framework and harnesses to post-train an in-house model on company-specific workflows and tools.
- •Developer coding plans
Give engineers a $6-$100 monthly Token Plan that backs Claude Code, OpenCode or OpenClaw sessions with MiMo models.
Ideal For
Best For
- ✓Enterprises that need a near-frontier model they can run on their own infrastructure under a permissive MIT licence for data-control reasons
- ✓High-volume coding-agent and agentic workloads where API cost dominates, using Flash at $0.14/$0.28 per million input/output tokens
- ✓Long-document and multimodal analysis that needs a 1M-token context across text, images, audio and video in one model
- ✓Teams post-training their own agents with Xiaomi's released RL framework and 7,000+ task environments
Not Ideal For
- ✗Organizations whose policies bar sending data to a China-based API provider; they can only use MiMo by self-hosting the weights
- ✗Teams without multi-node GPU clusters who want to self-host Pro; Flash or a third-party host is more realistic
- ✗Buyers who need contractual SLAs, SOC 2/ISO attestations or data-residency guarantees from the vendor, none of which Xiaomi publishes for the MiMo API
- ✗Workloads that require the very best coding-agent performance, where Claude Opus 5 still leads on DeepSWE, ProgramBench and Terminal Bench
Deployment
Market Analysis
Pros
- ✓MIT licence permits commercial use, fine-tuning and fully on-premises deployment
- ✓Top open-weights intelligence score at launch (Artificial Analysis 46) at a fraction of proprietary API prices
- ✓1M-token multimodal context and 128K output suit long-horizon agent work
- ✓Training is unusually transparent, with a technical report, RL environments and a live post-training dashboard that Hacker News commenters praised
Cons
- ✗Self-hosting Pro needs a multi-node H100/H200-class cluster; Hacker News commenters note even 128GB local machines are out of reach for these models
- ✗China-origin data and geopolitical concerns rule out the hosted API for some enterprises, a recurring theme in the launch discussion
- ✗Agent benchmark results are largely vendor-run, it still trails Claude Opus 5 on key coding benchmarks, and HN commenters questioned benchmark overfitting
- ✗Verbose: Artificial Analysis recorded 140M output tokens to run its index, and Decrypt's review of the earlier V2-Pro flagged token burn and weak frontier math
- ✗No published SLA, security certifications or data-residency commitments for the MiMo API
Pricing
Open weights (self-host)
$0
- ✓MIT licence, commercial use allowed
- ✓Pro, Flash and 9B distilled weights on Hugging Face
- ✓BF16, FP8 and quantized formats
- ✓Bring your own GPUs
API: MiMo-V2.6-Pro
$0.435 / 1M input, $0.87 / 1M output tokens
- ✓Cached input $0.0036 per 1M tokens
- ✓1M-token context
- ✓Multimodal input
API: MiMo-V2.6-Flash
$0.14 / 1M input, $0.28 / 1M output tokens
- ✓Cached input $0.0028 per 1M tokens
- ✓1M-token context
- ✓Roughly one-third the cost of Pro
Token Plan Lite
From $6/mo
- ✓4.1B credits per month
- ✓For coding tools such as Claude Code, OpenCode, OpenClaw
Token Plan Standard
From $16/mo
- ✓11B credits per month
Token Plan Pro
From $50/mo
- ✓38B credits per month
Token Plan Max
From $100/mo
- ✓82B credits per month
The weights are free under MIT for commercial use, so the real self-hosting cost is GPUs: the Pro model card recommends a two-node H100/H200-class deployment. Xiaomi's API charges $0.435/$0.87 per million input/output tokens for Pro ($0.0036 cached) and $0.14/$0.28 for Flash; Token Plan subscriptions cost $6, $16, $50 and $100 a month for credit bundles, and service is suspended when a monthly quota runs out. No enterprise tier, SLA or data-residency option is published. Artificial Analysis notes the model is verbose, generating 140M output tokens to run its index, which erodes some of the per-token saving.
Security & Compliance
Connect
Sources
This page was written from 10 sources, 10 on domains other than mimo.xiaomi.com.
- 1.venturebeat.com — better than deepseek xiaomis mimo v2 6 pro debuts as the top
- 2.huggingface.co — MiMo V2.6 Pro RL
- 3.huggingface.co — XiaomiMiMo
- 4.artificialanalysis.ai — mimo v2 6 pro
- 5.mimo.mi.com — token plan
- 6.en.wikipedia.org — Xiaomi MiMo
- 7.en.wikipedia.org — Xiaomi
- 8.hn.algolia.com — 49792730
- 9.decrypt.co — xiaomi mimo v2 pro review so good mistaken deepseek v4
- 10.mycodingplan.com — xiaomi mimo token plan
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
TypeSafe Jev
A 'System One' decision model that returns typed choices, scores and calibrated probabilities instead of text, for agent routing and classification
PrismML Bonsai
Open-weight 1-bit and ternary LLMs that run 27B-class AI on laptops and phones
Sakana AI Fugu Max
Multi-agent orchestration behind one API — frontier-grade results at $2 per million input tokens
Arcee Trinity
US-built open-weight model family, from on-device Trinity Nano to the 400B-parameter Trinity Large, that you can run on your own infrastructure