A multi-institution empirical study (Shanghai AI Laboratory with Oxford, USTC, NUS, University of Georgia and others; senior author Lei Bai) of how model scale, data volume and compute budget interact in RL post-training for mathematical reasoning, run across the full Qwen2.5 dense series from 0.5B to 72B — the model-size axis of the RL-scaling literature, where ScaleRL covers recipe/compute curves and IsoCompute covers sampling allocation.

Headline finding: larger models are consistently more learning-efficient in RL on both compute and data metrics, with the paper characterizing the model-scale/data/compute relationships as empirical scaling behaviors rather than a single fitted law. Accepted to ACL 2026 Main.

Paper

Venue ACL 2026
Authors: Zelin Tan · Hejia Geng · Xiaohang Yu · Mulei Zhang · Guancheng Wan · Yifan Zhou · Qiang He · Xiangyuan Xue · Heng Zhou · Yutao Fan · Zhongzhi Li · Zaibin Zhang · Guibin Zhang · Chen Zhang
rlrl-scalingscaling-lawspost-trainingresearch

Related