RUC's entropy-based adaptive-rollout RL for tool-using agents (ICLR 2026; 1.1K★, with Kuaishou): allocate more exploration where the policy is uncertain, improving multi-turn agentic RL — heavily adopted in follow-on agent-training work. Part of RUC's deep-research / agent line (Search-o1, WebThinker, DeepAgent).

Paper

post-trainingagentsresearch