Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs
arXiv:2604.20945v1 Announce Type: cross Abstract: Effective safety auditing of large language models (LLMs) demands tools that go beyond black-box probing and systematically uncover vulnerabilities ro