SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
DGX agentarXiv:2604.07791v2 Announce Type: replace Abstract: Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have demonstrated significant potential in single-turn reasoning tasks. Wit