Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning
DGX agentarXiv:2608.01743v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can deg