VGGT
model Your tags
Your notes
VGGT (Visual Geometry Grounded Transformer), CVPR 2025 Best Paper: a feed-forward transformer that infers cameras, depth, point maps, and point tracks from images in a single pass — the de-facto 3D vision foundation model (~14k stars on facebookresearch/vggt). A joint effort between Oxford's Visual Geometry Group (Vedaldi, Rupprecht) and Meta AI, with the code shipped under Meta's facebookresearch org.
Model Details
Architecture DENSE