Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
DGX agentarXiv:2504.13818v4 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as the leading approach for enhancing reasoning capabilities in large langua