Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
arXiv:2604.18487v1 Announce Type: new Abstract: The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting fro