124B open-weights multimodal model (123B multimodal decoder + 1B vision encoder) "built on top of Mistral Large 2, i.e., Mistral-Large-Instruct-2407" — extending the text flagship with frontier-class image understanding without compromising text performance. 128K context fits at least 30 high-resolution images.

At release it led MathVista (CoT: 69.4, ahead of GPT-4o and Gemini-1.5 Pro), DocVQA (ANLS: 93.3), VQAv2 (80.9), and MM-MT-Bench (7.4, ahead of Claude 3.5 Sonnet), and was the best open-weights model on the LMSys Vision Arena by ~50 ELO. Mistral Research License (commercial license required for production use).

Model Details

Architecture DENSE
Parameters 124B
Context window 128,000
AA Intelligence 8
License Mistral Research License
Base model mistral-large-2
multimodalvisionopen-weight

Related