OpenAI benchmark of 750 life-science research tasks authored by 173 PhD scientists, graded against 19,020 rubric criteria across 7 workflow types (literature synthesis, experimental design, data analysis, artifact interpretation, and more). The best model at release scores 36.1%, with artifact interpretation the largest bottleneck — OpenAI's counterpart to the science-workflow evals emerging across frontier labs.

Evaluation Details

Questions 750
evalsciencebiology

Related