Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting
DGX agentarXiv:2601.02670v2 Announce Type: replace Abstract: We introduce self-jailbreaking, a threat model in which an aligned LLM guides its own compromise. Unlike most jailbreak techniques, which oft