LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards
DGX agentarXiv:2605.31584v1 Announce Type: cross Abstract: Long-context reasoning remains a central challenge for large language models, which often fail to locate and integrate key information in extensive di