Environment Evolution for Terminal Agents
paper Your tags
Your notes
Tencent Hunyuan (twelve authors, first author HKUST-affiliated) evolves verifiable terminal environments off-policy on a generation-scheduled difficulty curriculum, so that the environment distribution keeps pace with the agent being trained; validated on Qwen models. The environment axis of RL scaling, alongside Qwen's Terminal-Universe.