1. Overview: A Hardware Giant's AI Bet
On April 28, 2026, Xiaomi officially open-sourced the MiMo-V2.5 series β and in doing so, disrupted the global AI model landscape once again. A company best known for budget smartphones and electric vehicles just released a trillion-parameter language model under the permissive MIT license, making the weights freely available to developers worldwide.
The series comprises two distinct models: MiMo-V2.5 (310B total / 15B active, native omnimodal) and MiMo-V2.5-Pro (1.02T total / 42B active, optimized for complex agentic and software engineering tasks). Both support a 1 million token context window, employ a sparse Mixture-of-Experts (MoE) architecture with a hybrid attention design, and are available on the official API platform and Hugging Face simultaneously.
On March 11, 2026, a nameless model appeared on OpenRouter with no documentation: one trillion parameters, 1M context, zero cost. Within seven days it processed over 1 trillion tokens total and topped OpenRouter daily usage charts. The community was convinced it was DeepSeek V4. On March 18, Xiaomi MiMo head Luo Fuli revealed the truth: Hunter Alpha was an early internal test build of MiMo V2-Pro. Xiaomi shares jumped 5.8%.
2. Architecture Deep Dive
2.1 Hybrid Attention
V2.5-Pro interleaves Sliding Window Attention (SWA) and Global Attention (GA) at a 6:1 ratio with a 128-token sliding window. Combined with a learnable attention sink bias, this reduces KV-cache storage by nearly 7Γ while sustaining coherent reasoning across the full 1M-token context window.
2.2 Native Multi-Token Prediction (MTP)
Unlike traditional speculative decoding bolted on at inference, MiMo V2.5-Pro integrates three lightweight MTP modules using dense FFNs directly into training and inference. The result: the model predicts multiple future tokens in parallel, pushing output throughput to roughly 65.8 tokens/second at standard load.
2.3 Three-Stage Post-Training (V2.5-Pro)
Following the MiMo V2-Flash paradigm: general alignment β domain-specialized RL training (math, safety, code, agentic tool use, each with dedicated reward functions) β Multi-teacher On-Policy Distillation (MOPD), where the student model iteratively receives token-level guidance from all expert teachers simultaneously.
3. Benchmark Performance
All data sourced from Hugging Face official model cards and Artificial Analysis evaluations, collected AprilβMay 2026.
The headline numbers tell a clear story: MiMo V2.5-Pro is elite-tier at coding. SWE-bench Verified at 78.9% and SWE-bench Pro at 57.2% (real-world startup codebase bug fixes) place it firmly in the global top tier. GSM8K at 99.6% confirms strong mathematical reasoning, while HLE at 48% reveals a gap in cross-domain expert-level knowledge questions. Artificial Analysis gives it an overall Intelligence Index score of 54, well above the 33-point median for comparable open-weight models, though it is notably verbose: 92M output tokens generated during evaluation vs. the 41M average.
4. MiMo-V2.5: The Omnimodal Specialist
Where V2.5-Pro focuses on code and long-horizon text reasoning, MiMo-V2.5 takes a different path: it is a native omnimodal model that processes text, images, video, and audio within a single unified architecture β no stitched-together separate models.
The standout numbers here are Video-MME (87.7) and CharXiv RQ (81.0 β nearly matching GPT-5.4's 81.2 on academic chart understanding). In video understanding, V2.5 is competitive with Gemini 3 Pro. ClawEval Multimodal at 23.8 is the clearest weak point, indicating that native multimodal agentic execution still lags the pure-text Pro model's capabilities.
5. Agentic Capability: Engineering at Scale
The most compelling evidence for V2.5-Pro's capabilities is not benchmarks β it is what the model does when left to run autonomously on real engineering problems.
Based on a Peking University compiler principles course project: implement a complete SysY compiler in Rust including lexer, parser, AST, Koopa IR codegen, RISC-V assembly backend, and performance optimization. A PKU CS student typically needs several weeks. MiMo V2.5-Pro completed it in 4.3 hours / 672 tool calls / 233 of 233 tests passed.
From simple prompts, the model autonomously built a desktop video editor with multi-track timeline, clip trimming, cross-fades, audio mixing, and export pipeline. Final output: 8,192 lines of code, 11.5 hours, 1,868 tool calls.
Graduate-level analog circuit task: design and optimize a complete FVF-LDO (Flipped-Voltage-Follower low-dropout regulator) in TSMC 180nm CMOS. Six metrics must be met simultaneously. Wired into an ngspice simulation loop, V2.5-Pro converged in about one hour, with four key metrics improved by an order of magnitude over its initial design.
6. Ecosystem & Deployment
The V2.5 series launched with immediate Day-0 support from SGLang and vLLM β the two most widely deployed open-source inference engines. Hardware partnerships with AWS, AMD, T-HEAD, and Enflame ensure compatibility from cloud H100s to Chinese domestic accelerators. Agentic framework integrations include OpenCode Go, OpenClaw, KiloCode, Blackbox, and Cline. Developers on Claude Code can route to MiMo by pointing ANTHROPIC_BASE_URL to Xiaomi's endpoint. OpenRouter access is available at xiaomi/mimo-v2.5-pro.
7. Pricing: The Cost Disruption Story
| Model | Input ($/M) | Output ($/M) | Context | Open-Source |
|---|---|---|---|---|
| MiMo-V2.5 | $0.40 | $2.00 | 1M | β MIT |
| MiMo-V2.5-Pro | $1.00 | $3.00 | 1M (2Γ above 256K) | β MIT |
| Claude Opus 4.6 | $5.00 | $15.00 | 200K | β |
| GPT-5.4 | $5.00 | $30.00 | 128K | β |
| Gemini 3.1 Pro | $2.50 | $10.00 | 1M | β |
Beyond raw pricing, Xiaomi's token efficiency advantage is significant: V2.5-Pro uses 40β60% fewer tokens than comparable models to complete the same agentic task, meaning the effective cost gap is wider than the sticker price suggests. For context windows above 256K tokens, the price doubles to $2.00 input / $6.00 output β a notable consideration for long-document workflows. A Token Plan subscription is also available at four tiers ranging from $63.36/year (720M credits) to $1,056/year (19.2B credits).
8. Competitive Radar
The radar tells the story directly: MiMo V2.5-Pro's edge is in coding and math reasoning. Its agentic performance is competitive with Claude Opus 4.6. The multimodal axis is zero β V2.5-Pro is text-only, by design. That's precisely why the two models in this series complement each other rather than compete.
9. Strengths & Weaknesses
β Strengths
- Elite coding agent: SWE-bench Verified 78.9%, SWE-bench Pro 57.2% β top-tier globally for open-source models
- Exceptional price-to-performance: $1/M input tokens vs. $5 for Claude Opus 4.6 β a 5Γ cost advantage at comparable agentic quality
- Token efficiency: 40β60% fewer tokens per task than leading competitors, making the actual cost gap even larger
- True 1M context: GraphWalks shows stable performance at 512K and partial retention at 1M β a genuine improvement over V2-Pro
- MIT license: No authorization required, freely deployable commercially and self-hostable
- Day-0 ecosystem support: SGLang, vLLM, OpenRouter, and major agentic frameworks at launch
β οΈ Weaknesses
- Verbose output: 92M tokens generated during Artificial Analysis eval vs. 41M average β high output costs at scale
- Long thinking time: Some users report excessive "thinking time" for complex tasks, hurting real-time UX
- Limited self-correction on subtle bugs: Without explicit error feedback, bug localization on complex issues lags DeepSeek V4 Pro
- Extended context pricing doubles above 256K: Significant cost jump for true long-document use cases
- V2.5-Pro is text-only: No image/video/audio β need V2.5 base for multimodal tasks
10. Score Summary
| Dimension | MiMo V2.5 | MiMo V2.5-Pro |
|---|---|---|
| Coding / Engineering | ββββ | βββββ |
| Multimodal Understanding | ββββ | N/A |
| Price Competitiveness | βββββ | βββββ |
| Openness | βββββ | βββββ |
| Knowledge Breadth | βββ | βββ |
| Overall | 4.2 / 5 | 4.4 / 5 |
11. Who Should Use MiMo V2.5?
- Code agent developers: V2.5-Pro is among the best open-source models for SWE-bench-class tasks today. Try it first.
- Multimodal app builders: V2.5's native image/video/audio understanding at $0.40/M input is hard to beat for the price.
- Cost-sensitive enterprises: 5β12Γ cheaper than closed-source equivalents, combined with 40β60% token efficiency advantage.
- Self-hosting teams: MIT license + SGLang/vLLM + domestic accelerator support lowers the deployment barrier significantly.
- Not ideal for: Real-time interactions where thinking latency matters, or multimodal agentic tasks requiring V2.5-Pro-class reasoning (use V2.5 + V2.5-Pro in combination instead).
Verdict: MiMo V2.5-Pro is the most cost-competitive elite coding and agentic model available in open-source form as of May 2026. At one-fifth the price of Claude Opus 4.6, it matches or exceeds frontier benchmarks on software engineering tasks and ships with a fully permissive MIT license. Xiaomi β a phone company β just made a strong case for being taken seriously as a foundational AI lab.
This review was generated with the assistance of AI technology.