MiMo-V2.5
model Your tags
Your notes
Xiaomi's omnimodal sparse-MoE model (310B total / 15B active), distinct from the MiMo-V2.5-Pro flagship: hybrid attention (sliding-window + global), 1M context, and native text/image/video/audio via a 729M ViT and 261M audio encoder, trained on ~48T tokens with FP8 and agentic-RL post-training. SWE-Bench Pro 56.1. Shipped with a faster MiMo-V2-Flash sibling.
Model Details
Architecture MOE
Parameters 310B
Active params 15B
Context window 1,000,000
AA Intelligence 37