MiMo-V2.6-Flash
modelYour notes
The efficiency-balanced sibling of MiMo-V2.6-Pro, released the same day as open weights under MIT. 309B total / 15B active sparse MoE: 48 layers (39 sliding-window with a 128-token window, 9 global), hidden size 4096, 256 routed experts with 8 active and no shared experts, the same 5-layer DFlash-style MTP drafter, 681M MiMo ViT, and 308M + 127M audio encoders as Pro. Native text, image, video, and audio input with a 1M context and 128K max output. Pretrained on 48T tokens (26T text, 22T omni), more than the Pro's 30T, with the same AdamW recipe, 32K to 256K context extension, and MXFP4 quantization-aware mid-training. It went through the same single mixed RL run and groupwise agentic grading as Pro, at $0.9M of RL post-training versus Pro's $2.6M, with the run livestreamed on Xiaomi's public dashboard.
Reported scores sit close behind Pro on agentic work and ahead of it on one cyber benchmark: DeepSWE v1.1 67.9 (Pro 71.9), Terminal-Bench 4.0 28.8, Terminal-Bench 2.1 87.6, OSWorld-Verified 80.8, Toolathlon-Verified 73.6, AutomationBench 52.3, JobBench 61.2, CyberGym 95.1 (Pro 94.0), SEC Bench Pro 47.5, MiMo VisualCoding 71.5. Xiaomi positions it as the Pareto point for high-frequency calls: $0.14/M input and $0.28/M output (cache hits $0.0028), a third of Pro's price. Successor to the 310B/15B MiMo-V2.5. Not yet scored by Artificial Analysis at filing.
Model Details
Benchmark Scores
| Benchmark | Score | Mode |
|---|---|---|
| DeepSWE v1.1 | 67.9 | — |
| Terminal-Bench 4.0 | 28.8 | — |
| OSWorld-Verified | 80.8 | — |
| Toolathlon-Verified | 73.6 | — |
| CyberGym | 95.1 | — |