Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models
DGX agentarXiv:2606.30899v1 Announce Type: cross Abstract: Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified malicious beh