LLaVA-OneVision-2
model Your tags
Your notes
"Towards next-generation perceptual intelligence": a fully open 8B LMM unifying image, long-video, and spatial understanding, with the entire pipeline released — data recipes, encoders, training logs, and the lean lmms-engine unified multimodal training stack that now backs LMMs-Lab releases.
NTU-led LMMs-Lab release (Ziwei Liu, Bo Li), continuing the open-everything trajectory of LLaVA-OneVision-1.5; sibling efforts include OpenMMReasoner (CVPR 2026, transparent SFT+RL multimodal-reasoning recipe) and the LongVT "thinking with long videos" line.
Model Details
Variants
| Name | Parameters | Notes |
|---|---|---|
| LLaVA-OneVision-2-8B | — | fully open pipeline (data, encoders, logs) |