Rollout-Level Advantage-Prioritized Experience Replay for GRPO
DGX agentarXiv:2606.04560v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards with GRPO is a standard approach for post-training reasoning LLMs. It remains sample inefficient. Each