Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
DGX agentarXiv:2506.07468v4 Announce Type: replace-cross Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders