A long-form movie and TV benchmark for evaluating long-video multimodal LLMs, from the Mohamed Elhoseiny / Vision-CAIR group. It stresses hour-plus temporal reasoning and has become a standard yardstick within the long-video community, used to evaluate Goldfish, LongVU, and Tempo.

Evaluation Details

Domains 2
Scoring accuracy
Domains: video, multimodal
evaluationbenchmarkvideomultimodal

Related