A verified version of SWE-Bench Pro, the repository-level coding benchmark that frontier labs now report alongside SWE-bench Verified. The authors (ECNU and Fudan with Shanghai AI Lab's OpenCompass team) find the original undermined by reward hacking, through leaked gold solutions and hidden evaluation information, and by task-quality defects such as misleading problem statements and badly scoped tests. The verified set adds anti-hacking safeguards that close the leakage channels without breaking normal agent behaviour and refines the affected tasks (102 corrected), so scores reflect coding ability rather than exploits. Shipped in AgentCompass and as a HuggingFace dataset.

Paper

evalbenchmarkcodingagents

Related