Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs
DGX agentarXiv:2606.07963v1 Announce Type: new Abstract: Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific trigg