Staleness-Learning Rate Scaling Laws for Asynchronous RLHF
DGX agentarXiv:2607.01083v1 Announce Type: cross Abstract: High-throughput RLHF systems often decouple rollout generation from policy optimization, leading to the use of stale rollouts during learner updates.