"Is Your Code Generated by ChatGPT Really Correct?" (NeurIPS 2023): HumanEval+ and MBPP+ add 80×/35× more tests via automated generation, exposing pass@1 drops of up to ~29% on results reported against the originals. Pre-window legacy seed, but still the default rigorous code eval with a continuously maintained live leaderboard from Lingming Zhang's group.

Paper

benchmarkcodingevaluation

Related