NVIDIA's case that the sequential axis of RL compute genuinely expands capability rather than just sharpening it — the direct counterpoint to the pass@k limits result. ProRL stabilizes RL over thousands of steps via KL divergence control, reference-policy resetting, and a diverse task suite, and shows the trained models beating their base across pass@k evaluations including cases where the base fails at every sampled attempt — evidence of novel reasoning strategies inaccessible to the base model under any sampling budget. Boundary expansion correlates with the base model's initial weakness on a task.

Released with open weights (Nemotron-Research-Reasoning-Qwen-1.5B). In the RL-scaling literature ProRL anchors the training-steps axis: its own plateau after thousands of steps is the launch point for BroRL's rollout-width scaling by the same group.

Paper

Authors: Mingjie Liu · Shizhe Diao · Ximing Lu · Jian Hu · Xin Dong · Yejin Choi · Jan Kautz · Yi Dong
rlrl-scalingpost-trainingreasoningresearch

Related