Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces
DGX agentarXiv:2605.02801v1 Announce Type: new Abstract: As large language model (LLM) agents evolve from isolated tool users into coordinated teams, reinforcement learning (RL) must optimize not only individu