"You Don't Need That Much Data to Train a Search Agent via RL": decouples the searcher from a frozen generator and optimizes a Gain-Beyond-RAG reward, matching or beating full search-RL pipelines with only ~2.4K training examples — orders of magnitude less than Search-R1-style end-to-end training. The data-efficiency follow-up in Jiawei Han's agentic-search line.

Paper

agenticefficiencytraining

Related