The first open large multimodal model built as a general evaluator / reward model, from NTU's LMMs-Lab with ByteDance — it judges other vision-language models' outputs and provides reward signals for preference optimization, and seeded a follow-on line (LLaVA-Critic-R1). Complements the filed LMMs-Eval harness and LLaVA-OneVision line.

Paper

multimodalevalopen-weight

Related