ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
DGX agentarXiv:2601.06487v3 Announce Type: replace-cross Abstract: Reinforcement learning has substantially improved the performance of LLM agents on tasks with verifiable outcomes, but it still struggles on o