Meta's large-scale answer to the missing pretraining-style scaling methodology for RL: a 400,000+ GPU-hour systematic study (work done at Meta; first author Devvrit Khatri of UT Austin, with UCL, Berkeley, Harvard and Periodic Labs co-affiliations, senior author Rishabh Agarwal) that fits sigmoidal compute–performance curves to RL training — pass rate vs log-compute, rather than pretraining's power laws — and ablates common recipe choices against two separated quantities: asymptotic performance and compute efficiency.

Three findings organize the study: recipes differ in their asymptote, not just speed; details like loss aggregation, normalization, curriculum and off-policy algorithm mostly modulate efficiency without moving the asymptote; and stable recipes follow predictable trajectories, so large-run performance can be extrapolated from small runs. The distilled best-practice recipe, ScaleRL, is validated by predicting the validation-performance curve of a single RL run scaled to 100,000 GPU-hours. Published at ICLR 2026; the founding entry of the RL-compute-scaling literature that IsoCompute extends into allocation rules.

Paper

Venue ICLR 2026
Authors: Devvrit Khatri · Lovish Madaan · Rishabh Tiwari · Rachit Bansal · Sai Surya Duvvuri · Manzil Zaheer · Inderjit S. Dhillon · David Brandfonbrener · Rishabh Agarwal
rlrl-scalingscaling-lawspost-trainingresearch

Related