Video-MMMU
eval Your tags
Your notes
Measures knowledge acquisition from professional videos: 300 expert-level videos and 900 human-annotated questions across six disciplines, structured along three cognitive stages — Perception, Comprehension, Adaptation — with a Δknowledge metric quantifying how much a model improves after "watching" the video. Revealed steep performance decline as cognitive demand rises and a wide human-model gap; adopted into frontier VLM evaluation suites (reported by Gemini and Qwen-VL releases).
LMMs-Lab / NTU (Ziwei Liu's orbit), evaluated via LMMs-Eval.
Paper
Evaluation Details
Questions 900
Domains 6
Domains: art, business, science, medicine, humanities, engineering