Shifting the Gradient: Understanding How Defensive Training Methods Protect Language Model Integrity
DGX agentarXiv:2604.16423v1 Announce Type: new Abstract: Defensive training methods such as positive preventative steering (PPS) and inoculation prompting (IP) offer surprising results through seemingly simila