ScienceArena: Benchmarking LLMs on Latest Scientific Olympiads
paper Your tags
Your notes
Qiyuan Tech with Tsinghua, HKU and Peking (the Harness-Bench collaboration pattern) answers benchmark saturation and contamination with an olympiad-style eval built from thirteen recent public science competitions — IPhO and IChO 2025–26, IBO 2023, USAPhO 2026, USNCO 2025 — whose open-ended multi-step problems carry process-credit rubrics for step-level scoring, matching the reasoning-process-eval criterion the index tracks.
An expert-audited digitization pipeline converts official exams, figures, solutions and rubrics into structured items verified by olympiad medalists, and LLM-as-judge scoring is calibrated against medalist ground truth.