Scalable On-Policy Reinforcement Learning via Adaptive Batch Scaling
DGX agentarXiv:2605.21557v1 Announce Type: cross Abstract: Conventional wisdom holds that large-batch training is fundamentally incompatible with Reinforcement Learning (RL) - beyond a modest threshold, increa