Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
DGX agentarXiv:2605.07114v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language model