Effective Reinforcement Learning for Agentic Search by Recycling Zero-Variance Queries During Training
DGX agentarXiv:2606.10709v1 Announce Type: cross Abstract: The use of GRPO-style algorithms has become the standard strategy for training LLM search agents under outcome-only rewards. With these algorithms, a