VLM-3R
paper Your tags
Your notes
Vision-language model augmented with instruction-aligned monocular 3D-reconstruction tokens (CVPR 2026): fuses geometry from video directly into the VLM for spatial reasoning without depth sensors, and introduces the VSTI benchmark for visual spatial-temporal intelligence (427★). VITA group work continuing under Zhangyang Wang's UT students while he is on leave at XTX Markets.