Measures knowledge acquisition from professional videos: 300 expert-level videos and 900 human-annotated questions across six disciplines, structured along three cognitive stages — Perception, Comprehension, Adaptation — with a Δknowledge metric quantifying how much a model improves after "watching" the video. Revealed steep performance decline as cognitive demand rises and a wide human-model gap; adopted into frontier VLM evaluation suites (reported by Gemini and Qwen-VL releases).

LMMs-Lab / NTU (Ziwei Liu's orbit), evaluated via LMMs-Eval.

Paper

Evaluation Details

Questions 900
Domains 6
Domains: art, business, science, medicine, humanities, engineering
benchmarkevaluationvideomultimodal

Related