A Sober Look at Progress in LM Reasoning
paper Your tags
Your notes
COLM 2025, from the Bethge/evaluation-science line: much of the reported gain from RL fine-tuning for reasoning evaporates under standardized, seed- and decoding-controlled evaluation. A widely-cited rigor check that reshaped how RL-for-reasoning claims are reported — sibling to the cluster's other benchmark-science work.