From Base Rollouts to RL Reasoning (A Budgeted Search Perspective)
paper Your tags
Your notes
Findings of EMNLP 2026, from Fudan with Zhipu AI, Tsinghua, and the Shanghai Innovation Institute. A unified decoding framework asks whether RLVR creates new reasoning ability or reshapes the base model's rollout distribution, measured through pass@k and best-of-N under matched budgets. The mechanisms-and-limits axis of the RL-scaling literature.