The first open LMM family to transfer across single-image, multi-image, and video from one recipe: 0.5B/7B/72B models on Qwen2 LLM backbones with SigLIP vision and AnyRes, trained on the fully released OneVision data mixture. The 7B checkpoint still sees ~225K monthly HuggingFace downloads — among the most-used open multimodal models ever shipped from academia.

Joint NTU LMMs-Lab / ByteDance co-development, not sole NTU. Direct spawns from the same collaboration: LLaVA-Video (synthetic video-instruction data + 7B/72B models) and LLaVA-Critic (the first open multimodal judge/reward LMM, with the LLaVA-Critic-R1 follow-up). Succeeded by the fully open LLaVA-OneVision-1.5.

Model Details

Base model qwen2

Variants

Name Parameters Notes
llava-onevision-qwen2-0.5b-ov Qwen2-0.5B backbone
llava-onevision-qwen2-7b-ov Qwen2-7B backbone, ~225K monthly downloads
llava-onevision-qwen2-72b-ov Qwen2-72B backbone

Paper

multimodalvisionvideoopen-weight

Related