Safety
Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence
arXiv:2608.20820v1 Announce Type: new Abstract: Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robus
arXiv:2608.20820v1 Announce Type: new Abstract: Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to single-turn inputs; naive multi-turn composition yields bounds that degrade exponentially in the number of turns. We introduce Multi-Turn Certified Robustness (MTCR), a framework that models conversational safety via State-Adversarial MDPs and defines k-turn certified robustness as the worst-case safety probability across k adversarial turns. MTCR comprises: (i) compositional certification via embedding-space mode decomposition, yielding tighter certified lower bounds than naive multiplication; (ii) (alpha,eta)-safety persistence, improving the degradation rate from nderline{p}^{k} to eta^k (with eta > nderline{p}) and yielding interpretable horizon estimates; (iii) matching information-theoretic upper bounds establishing tightness; and (iv) a unified algorithm combining these results. Experiments on six LLMs under epsilon-bounded and Crescendo-style attacks confirm that empirical safety consistently exceeds the certified bounds.
Related
- Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models
- Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons
- D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting
- THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models
Source: arXiv cs.AI | 2026-08-24