Step-3-VL-10B
model paper Your tags
Your notes
"Compact giant" outperforming models 20x its size (including GPT-4o on certain benchmarks) via fully unfrozen perception-decoder training.
Outputs 2
Step-3-VL-10B
model"Compact giant" 10B vision-language model outperforming models 20x its size (including GPT-4o on certain benchmarks) via fully unfrozen perception-decoder training on 1.2T multimodal tokens; the decoder starts from Qwen3-8B. AA Intelligence Index v4.3: 8.
Architecture DENSE
AA Intelligence 8
Base model qwen3
Released Jan 20, 2026 on HuggingFace. Decoder initialized from Qwen3-8B (fully unfrozen 1.2T-token multimodal pretraining), so the 10B footprint is recorded in the description rather than as from-scratch scale.
Step3-VL-10B Technical Report
paperFocused on "Intrinsic Vision-Language Synergy."