Ling-3.0-flash-VL
modelYour notes
Ant Group's first native multimodal Ling model, built on Ling-3.0-flash: a 124B MoE that activates 5.5B parameters per token (512 experts, 8 active), takes images and video, and runs 131K natively with 256K via YaRN. A ViT encoder and two-layer MLP projector feed the 42-layer hybrid backbone, which alternates KDA and Gated MLA layers at 5:1; VideoRoPE encodes spatial position and temporal order so the model can localize events and answer questions over long video. Ant frames the release around agentic use, with vision carried through understanding, reasoning, planning, acting, and verification rather than treated as an input only. Thinking mode is on by default. MIT license; Ant publishes SGLang recipes for 4-GPU H200/H20 and Blackwell nodes at full 256K context. Announced September 9, 2026 with weights public the next day and free access on Ling Studio; Ant cites an Image-to-WebDev Arena result above GPT-5.4.
Artificial Analysis scored it within days: 25 on Intelligence Index v4.3, level with the text-only Ling-3.0-flash; Ant's own Terminal-Bench 2.1 number follows the AA protocol (Terminus 2 harness, three runs per task). Ant also shipped a finance-tuned sibling, Ling-3.0-flash-Fin, the day before.