sGPO: Trading Inference FLOPs for Training Efficiency in RLVR
DGX agentarXiv:2606.08854v1 Announce Type: cross Abstract: Standard Reinforcement Learning with Verifiable Rewards (RLVR) training allocates a fixed rollout budget to every query, without regard for what each