Model Releases

Safety Training May Persist Through Helpfulness Optimization in LLM Agents

arXiv:2603.02229v2 Announce Type: replace-cross Abstract: Safety post-training has been studied extensively in single-step 'chat' settings where safety typically refers to refusing harmful requests. W

DGX agentpaper
model-releasesarxiv-cs-cl

arXiv:2603.02229v2 Announce Type: replace-cross Abstract: Safety post-training has been studied extensively in single-step "chat" settings where safety typically refers to refusing harmful requests. We study an "agentic" (i.e., multi-step, tool-use) setting where safety refers to harmful actions directly taken by the LLM. We investigate the effects of using direct preference optimization (DPO) to optimize safety and/or helpfulness on the ToolEmu agentic benchmark. First, we find that safety training largely persists through subsequent helpfulness training. Second, we find a consistent negative linear correlation (R^2 = 0.77) between safety and helpfulness when considering all training configurations together. Even post-training on both metrics simultaneously simply results in another point on the same trend line rather than yielding a "best of both worlds" strategy, despite the presence of such strategies in our dataset. Overall, our findings underscore the need for a better understanding of post-training.

Source: arXiv cs.CL | 2026-08-25

Loading related sources…