"Compact giant" outperforming models 20x its size (including GPT-4o on certain benchmarks) via fully unfrozen perception-decoder training.

Outputs 2

Step-3-VL-10B

model

"Compact giant" 10B vision-language model outperforming models 20x its size (including GPT-4o on certain benchmarks) via fully unfrozen perception-decoder training on 1.2T multimodal tokens; the decoder starts from Qwen3-8B. AA Intelligence Index v4.3: 8.

Architecture DENSE
AA Intelligence 8
Base model qwen3

Released Jan 20, 2026 on HuggingFace. Decoder initialized from Qwen3-8B (fully unfrozen 1.2T-token multimodal pretraining), so the 10B footprint is recorded in the description rather than as from-scratch scale.

Step3-VL-10B Technical Report

paper

Focused on "Intrinsic Vision-Language Synergy."

multimodalefficiencyopen-weight