Quick answer: MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of Xiaomi's MiMo-V2.6 series: a 309B-parameter sparse MoE (15B active) with text, image, video and audio input and a 1M-token context, released September 21, 2026 under the MIT license. It reports 87.6 on Terminal-Bench 2.1, 67.9 on DeepSWE v1.1 and 73.6 on Toolathlon-Verified — close to its 1.02T-parameter sibling at roughly a third of the active parameters.
Where it leads: Near-flagship agentic coding and computer use for its size: OSWorld-Verified 80.8, Terminal-Bench 2.1 87.6 and DeepSWE v1.1 67.9.
Where it lags: Terminal-Bench 4.0 (28.8), ExploitBench (25.3) and SEC-Bench Pro (47.5) are well below the Pro checkpoint and the closed frontier models.
Best for: Cost-efficient self-hosted agents and long-context coding workloads that still need omnimodal input.
It shares the MiMo-V2.6 recipe with MiMo-V2.6-Pro-RL: one mixed asynchronous GRPO run across coding, general agents, vision and cybersecurity, groupwise agentic grading, and multi-teacher on-policy distillation. The 48-layer backbone interleaves 39 sliding-window layers with 9 global-attention layers, 256 routed experts (8 active), a 681M-parameter vision encoder, audio encoders and a 5-layer multi-token-prediction drafter.
Weights are open (FP8, MIT license) and the model is available via the Xiaomi MiMo API platform, AI Studio, MiMo Code and OpenRouter.
| Field | Value |
|---|---|
| Organization | Xiaomi |
| Hugging Face | XiaomiMiMo/MiMo-V2.6-Flash-RL |
| Architecture | Sparse MoE, 309B total / 15B active |
| Context length | 1M tokens |
| Modalities | Text, image, video, audio |
| License | MIT |
| Release date | 2026-09-21 |
| Benchmark | Score | Source | Date |
|---|---|---|---|
| DeepSWE (v1.1) | 67.9 | Model card | 2026-09 |
| Terminal-Bench 4.0 | 28.8 | Model card | 2026-09 |
| Terminal-Bench 2.1 | 87.6 | Model card | 2026-09 |
| Toolathlon-Verified | 73.6 | Model card | 2026-09 |
| OSWorld-Verified | 80.8 | Model card | 2026-09 |
| JobBench | 61.2 | Model card | 2026-09 |
| ExploitBench | 25.3 | Model card | 2026-09 |
| SEC-Bench Pro | 47.5 | Model card | 2026-09 |
Self-reported by Xiaomi; not independent Benchgen measurements. Also reported but not added: ProgramBench 26.0 and Agents' Last Exam 27.6 (metric differs), AutomationBench v1.0.6 52.3 (version-mismatch precedent), CyberGym 95.1 and ExploitGym 6.0 (scale differs), and Xiaomi-internal MiMo Code Bench, MiMo Cyber Bench and MiMo VisualCoding.