Scaling Behaviors of LLM Reinforcement Learning Post-Training
paper Your tags
Your notes
A multi-institution empirical study (Shanghai AI Laboratory with Oxford, USTC, NUS, University of Georgia and others; senior author Lei Bai) of how model scale, data volume and compute budget interact in RL post-training for mathematical reasoning, run across the full Qwen2.5 dense series from 0.5B to 72B — the model-size axis of the RL-scaling literature, where ScaleRL covers recipe/compute curves and IsoCompute covers sampling allocation.
Headline finding: larger models are consistently more learning-efficient in RL on both compute and data metrics, with the paper characterizing the model-scale/data/compute relationships as empirical scaling behaviors rather than a single fitted law. Accepted to ACL 2026 Main.