Alignment Risks from Capability-Seeking RL Training
DGX agentarXiv:2602.12124v2 Announce Type: replace-cross Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capab