Benchgen
Models/xiaomi/

MiMo-V2.6-Flash-RL

DraftPublic

Model Details

MiMo-V2.6-Flash-RL

Organization License Params Context Released

Quick answer: MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of Xiaomi's MiMo-V2.6 series: a 309B-parameter sparse MoE (15B active) with text, image, video and audio input and a 1M-token context, released September 21, 2026 under the MIT license. It reports 87.6 on Terminal-Bench 2.1, 67.9 on DeepSWE v1.1 and 73.6 on Toolathlon-Verified — close to its 1.02T-parameter sibling at roughly a third of the active parameters.

At a Glance

Where it leads: Near-flagship agentic coding and computer use for its size: OSWorld-Verified 80.8, Terminal-Bench 2.1 87.6 and DeepSWE v1.1 67.9.

Where it lags: Terminal-Bench 4.0 (28.8), ExploitBench (25.3) and SEC-Bench Pro (47.5) are well below the Pro checkpoint and the closed frontier models.

Best for: Cost-efficient self-hosted agents and long-context coding workloads that still need omnimodal input.

What MiMo-V2.6-Flash-RL Is

It shares the MiMo-V2.6 recipe with MiMo-V2.6-Pro-RL: one mixed asynchronous GRPO run across coding, general agents, vision and cybersecurity, groupwise agentic grading, and multi-teacher on-policy distillation. The 48-layer backbone interleaves 39 sliding-window layers with 9 global-attention layers, 256 routed experts (8 active), a 681M-parameter vision encoder, audio encoders and a 5-layer multi-token-prediction drafter.

Weights are open (FP8, MIT license) and the model is available via the Xiaomi MiMo API platform, AI Studio, MiMo Code and OpenRouter.

Specifications

FieldValue
OrganizationXiaomi
Hugging FaceXiaomiMiMo/MiMo-V2.6-Flash-RL
ArchitectureSparse MoE, 309B total / 15B active
Context length1M tokens
ModalitiesText, image, video, audio
LicenseMIT
Release date2026-09-21

Public Benchmark Scores

Self-reported by Xiaomi; not independent Benchgen measurements. Also reported but not added: ProgramBench 26.0 and Agents' Last Exam 27.6 (metric differs), AutomationBench v1.0.6 52.3 (version-mismatch precedent), CyberGym 95.1 and ExploitGym 6.0 (scale differs), and Xiaomi-internal MiMo Code Bench, MiMo Cyber Bench and MiMo VisualCoding.

Frequently Asked Questions

What is MiMo-V2.6-Flash-RL? Xiaomi's efficiency-balanced MiMo-V2.6 checkpoint: a 309B-parameter omnimodal MoE (15B active) with a 1M-token context, released September 21, 2026 under the MIT license.
How is it different from MiMo-V2.6-Pro-RL? Pro-RL is the 1.02T-parameter (42B active) flagship; Flash-RL is about a third of the size with similar scores on most agentic benchmarks but a clear gap on cybersecurity tasks.
Is MiMo-V2.6-Flash-RL open weights? Yes, MIT-licensed weights are on Hugging Face, and it is also served through the Xiaomi MiMo API platform and OpenRouter.