CANOPY (Outcome-Only RL for Long-Horizon Agents)
paper Your tags
Your notes
"Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents." Alibaba Research diagnoses two failure modes of group-relative outcome-only RL on long interactive tasks, signal starvation and policy drift, and trains Qwen3-14B through environment interaction alone with the resulting recipe. Code under Apache 2.0.