"…from Small to Large LLMs" (EMNLP 2026). NUS (Jingtan Wang, Bryan Kian Hsiang Low) with A*STAR and SMART/MIT (Daniela Rus) study how to split a fixed annotation budget between SFT data and RL, and find that the near-optimal region widens with model size and transfers from small proxy models to large ones. The budget-allocation axis of RL scaling.

Paper

rlrl-scalingpost-training