The evaluation companion to Intern-S2's raw-page pretraining: a workflow-centred benchmark of 124 expert-authored, difficulty-screened questions in seven research-assistant capability groups and 19 subtasks across five scientific domains, each instantiated under four matched conditions (English or Chinese, image-first or markdown input) for 496 instances that require reasoning jointly over text, equations, figures, tables, code, and data while keeping the provenance of evidence. The best system scores 62.6 of 100. Released with the SciDocIR retrieval set and about 15K SFT and 8K RL training samples, and integrated into VLMEvalKit; authors include Jiaqi Wang, Yuhang Zang, and Dahua Lin.

Paper

evalbenchmarksciencemultimodal

Related