HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment
DGX agentarXiv:2607.00572v1 Announce Type: new Abstract: Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed a