An Empirical Study of Multi-Generation Sampling for Jailbreak Detection in Large Language Models
DGX agentarXiv:2604.18775v1 Announce Type: new Abstract: Detecting jailbreak behaviour in large language models remains challenging, particularly when strongly aligned models produce harmful outputs only rarel