The environment axis of RL scaling, from a UW-led team (Zhiyuan Zeng; senior authors Natasha Jaques, Simon Du, Yulia Tsvetkov) with Ai2 (Hamish Ivison), UIUC and LMSYS. RLVE replaces static RL datasets — whose learning signal vanishes when problems are too easy or too hard for the current policy — with verifiable environments that procedurally generate problems and adapt their difficulty distribution to the policy as training progresses, each with algorithmically checkable rewards.

Ships RLVE-Gym, a 400-environment suite built through manual environment engineering. Accepted to ICML 2026; IsoCompute cites it as the environment-scaling dimension of the LLM RL-scaling literature.

Paper

Venue ICML 2026
Authors: Zhiyuan Zeng · Hamish Ivison · Yiping Wang · Lifan Yuan · Shuyue Stella Li · Zhuorui Ye · Siting Li · Jacqueline He · Runlong Zhou · Tong Chen · Chenyang Zhao · Yulia Tsvetkov · Simon Shaolei Du · Natasha Jaques
rlrl-scalingpost-trainingenvironmentsresearch

Related