LLaVA-OneVision
model Your tags
Your notes
The first open LMM family to transfer across single-image, multi-image, and video from one recipe: 0.5B/7B/72B models on Qwen2 LLM backbones with SigLIP vision and AnyRes, trained on the fully released OneVision data mixture. The 7B checkpoint still sees ~225K monthly HuggingFace downloads — among the most-used open multimodal models ever shipped from academia.
Joint NTU LMMs-Lab / ByteDance co-development, not sole NTU. Direct spawns from the same collaboration: LLaVA-Video (synthetic video-instruction data + 7B/72B models) and LLaVA-Critic (the first open multimodal judge/reward LMM, with the LLaVA-Critic-R1 follow-up). Succeeded by the fully open LLaVA-OneVision-1.5.
Model Details
Base model qwen2
Variants
| Name | Parameters | Notes |
|---|---|---|
| llava-onevision-qwen2-0.5b-ov | — | Qwen2-0.5B backbone |
| llava-onevision-qwen2-7b-ov | — | Qwen2-7B backbone, ~225K monthly downloads |
| llava-onevision-qwen2-72b-ov | — | Qwen2-72B backbone |