Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
DGX agentarXiv:2606.12370v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a key component in modern large language models, yet the rollout stage remains the key bottleneck in RL trainin