Quick answer: MiMo-V2.6-Pro-RL is the flagship checkpoint of Xiaomi's MiMo-V2.6 series: a 1.02T-parameter sparse MoE (42B active) with text, image, video and audio input and a 1M-token context, released September 21, 2026 under the MIT license. It was trained with one mixed, large-scale RL run across coding, general agents, vision and cybersecurity, and reports 89.9 on Terminal-Bench 2.1, 71.9 on DeepSWE v1.1 and 76.9 on Toolathlon-Verified.
Where it leads: Open-weight agentic coding and tool use close to frontier closed models — within about 2 points of Claude Opus 5 on DeepSWE (74.0) and ahead of GPT-5.6 Sol on Terminal-Bench 2.1 (89.9 vs 88.8) in Xiaomi's table.
Where it lags: Terminal-Bench 4.0 (34.9 vs Opus 5's 49.0), ExploitBench (47.9 vs 70.0-78.5) and ExploitGym trail the closed frontier models.
Best for: Self-hosted, long-horizon coding and computer-use agents that need an MIT-licensed omnimodal model.
MiMo-V2.6 scales reinforcement learning toward self-improvement: fully asynchronous GRPO on very large batches (1,568 prompts x 16 rollouts per step), groupwise agentic grading that builds task-specific rubrics from contrasting rollouts, and multi-prefix multi-teacher on-policy distillation to extend capabilities to hard-to-verify tasks. The backbone interleaves sliding-window (60 layers) and global attention (10 layers) over 70 layers with 384 routed experts (8 active), plus a 681M-parameter vision encoder, audio encoders and a 5-layer multi-token-prediction drafter.
It is available as open weights (FP8) and through the Xiaomi MiMo API platform, AI Studio, MiMo Code and OpenRouter. A smaller sibling, MiMo-V2.6-Flash-RL, has 309B total and 15B active parameters.
| Field | Value |
|---|---|
| Organization | Xiaomi |
| Hugging Face | XiaomiMiMo/MiMo-V2.6-Pro-RL |
| Architecture | Sparse MoE, 1.02T total / 42B active |
| Context length | 1M tokens |
| Modalities | Text, image, video, audio |
| License | MIT |
| Release date | 2026-09-21 |
| Benchmark | Score | Source | Date |
|---|---|---|---|
| DeepSWE (v1.1) | 71.9 | Model card | 2026-09 |
| Terminal-Bench 4.0 | 34.9 | Model card | 2026-09 |
| Terminal-Bench 2.1 | 89.9 | Model card | 2026-09 |
| Toolathlon-Verified | 76.9 | Model card | 2026-09 |
| OSWorld-Verified | 82.0 | Model card | 2026-09 |
| JobBench | 62.0 | Model card | 2026-09 |
| GDPval-AA v2 (v2.1) | 1673 | Model card | 2026-09 |
| ExploitBench | 47.9 | Model card | 2026-09 |
| SEC-Bench Pro | 66.3 | Model card | 2026-09 |
Self-reported by Xiaomi; not independent Benchgen measurements. Also reported but not added: ProgramBench 26.5 and Agents' Last Exam 31.6 (metric differs from the Benchgen pages), AutomationBench v1.0.6 53.1 (version-mismatch precedent), CyberGym 94.0 and ExploitGym 17.8 (scale differs from the Benchgen pages), and Xiaomi-internal MiMo Code Bench, MiMo Cyber Bench and MiMo VisualCoding.