GeneBench-Pro
eval Your tags
Your notes
OpenAI benchmark of 129 research-level computational-biology problems. Because tasks are built from synthetic data generated from fully-known causal structure, grading is deterministic — no rubric judges. At release GPT-5.6 Sol Pro scores 31.5%; the best non-OpenAI model (Claude Opus 4.8) scores 16.0%. Ten questions released to HuggingFace and 50 shared with Artificial Analysis for independent scoring.
Evaluation Details
Questions 129