Xiaomi MiMo-V2.6 Pro Ships Under an MIT Licence
A trillion-parameter omnimodal model under MIT terms is unusual. The licence matters more than the benchmark position, and the parameter count matters more than both.
Xiaomi released MiMo-V2.6 Pro on 21 September 2026 under an MIT licence: 1.02 trillion total parameters, 42 billion active per token, a one million token context window, and text, image, audio and video input. A cheaper sibling, MiMo-V2.6 Flash, runs 310 billion total and 15 billion active. The benchmark position is respectable. The licence is the part worth reading twice.
Why the MIT Licence on MiMo-V2.6 Pro Matters
MIT is a genuine open source licence. It permits commercial use, modification, redistribution and private forks, with no field-of-use restriction, no revenue ceiling, and no clause that retires your rights when you compete with the people who trained the model.
That is worth stating plainly because the phrase open weights has been doing a lot of quiet work lately. Alibaba moved Qwen-Image-2.1 from Apache 2.0 to a non-commercial research licence, a change we covered when a model that looked open turned out to be research-only. Weights you can download and weights you can build a business on are different things, and the download button looks identical in both cases. If you are not sure which you have, the licence text is the only place that answers it.
Xiaomi also published more than 7,000 reinforcement learning task environments, an end-to-end RL framework, composable mini-harnesses, and a distilled 9B model, MiMo-V2.6-Distill-Qwen-9B. For most teams the 9B distillation is the practically useful artefact, in the way distillations usually are.
The Parameter Count Is Not the Hardware Requirement
A trillion parameters sounds unrunnable, and mostly it is, but the shape matters. MiMo-V2.6 Pro is a sparse mixture-of-experts model: 1.02 trillion parameters exist, 42 billion are active for any given token. Compute per token tracks the active count. Memory tracks the total, because every expert has to be resident somewhere the router can reach it.
That is the asymmetry people miss. Sparse models are cheap to run and expensive to host, which is the practical difference between a dense model and a mixture of experts. You do not need trillion-parameter compute. You do need somewhere to put a trillion parameters.
What It Costs Through the API
Xiaomi's own pricing, as reported at launch, is $0.435 per million uncached input tokens and $0.87 per million output for Pro, and $0.14 and $0.28 for Flash. VentureBeat reports MiMo-V2.6 Pro scoring 46 on Artificial Analysis' Intelligence Index, tied with the newly released Grok 4.7 and ahead of DeepSeek V4.1 Pro at 36.
Training cost is reported at roughly $2.62 million for Pro and $850,000 for Flash, completed in under six days. Treat those as the vendor's figures rather than audited ones, but the direction is consistent with everything else this year.
What You Could Realistically Run From Xiaomi's Release
Almost nobody reading this will self-host the Pro model, and the Flash variant at 310 billion total is not much friendlier. The artefact that matters for individuals is the 9B distillation, which is small enough for a workstation with a decent GPU and inherits the licence. That is the version to reach for if the appeal here is running something locally rather than having a fallback provider on paper.
It is also the version where the MIT terms do the most work, because a 9B model is realistic to fine-tune on your own data, and fine-tuning is exactly the activity restrictive licences tend to constrain.
What This Changes
For most builders, nothing immediately, and that is fine. An MIT-licensed model at this capability tier mainly changes your negotiating position: it is a credible fallback that no vendor can withdraw, price up, or restrict later. That is worth something even if you never deploy it, and it is the substantive part of the open versus closed question.
Sources: VentureBeat's launch coverage and the MiMo-V2.6-Pro model page on LLM-Stats, both accessed 22 September 2026.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


