Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs
arXiv:2608.09542v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent