Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
DGX agentarXiv:2509.09708v3 Announce Type: replace Abstract: Refusal on harmful prompts is a key safety behaviour in instruction-tuned large language models (LLMs), yet the internal causes of this behaviour re