Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
DGX agentarXiv:2512.05591v2 Announce Type: replace-cross Abstract: Large language model post-training relies on reinforcement learning to improve model capability and alignment quality. However, the off-policy