Faster Synchronous On-Policy RL via Straggler-Aware Group Sizing
arXiv:2606.02218v1 Announce Type: cross Abstract: Synchronous reinforcement learning methods such as Group Relative Policy Optimization (GRPO) provide stable and reproducible on-policy training, but t