The first widely-noted open "visual o1" (ICCV 2025; 2.1K★; with Tsinghua and Alibaba): structured multi-stage visual reasoning — summary, caption, reasoning, conclusion — that brought o1-style deliberate reasoning to vision-language models and was heavily cited across the multimodal-reasoning wave.

Paper

multimodalreasoningopen-weight