"On the Interplay between On-Policy Distillation and RLVR." NYU with UChicago, Waterloo, and Alberta: running on-policy distillation and then RLVR beats pure OPD, pure RLVR, and every joint fusion tried, an ordering the authors explain through pass@k coverage. A stage-aware result for post-training pipelines.

Paper

rlrl-scalingdistillation