Scaling Near-Optimal SFT-RL Annotation Budget Allocation
paper Your tags
Your notes
"…from Small to Large LLMs" (EMNLP 2026). NUS (Jingtan Wang, Bryan Kian Hsiang Low) with A*STAR and SMART/MIT (Daniela Rus) study how to split a fixed annotation budget between SFT data and RL, and find that the near-optimal region widens with model size and transfers from small proxy models to large ones. The budget-allocation axis of RL scaling.