Findings of EMNLP 2026, from Fudan with Zhipu AI, Tsinghua, and the Shanghai Innovation Institute. A unified decoding framework asks whether RLVR creates new reasoning ability or reshapes the base model's rollout distribution, measured through pass@k and best-of-N under matched budgets. The mechanisms-and-limits axis of the RL-scaling literature.

Paper

rlrl-scalingreasoning