Safety

Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

arXiv:2608.20820v1 Announce Type: new Abstract: Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robus

DGX agentpaper
safetyarxiv-cs-ai

arXiv:2608.20820v1 Announce Type: new Abstract: Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to single-turn inputs; naive multi-turn composition yields bounds that degrade exponentially in the number of turns. We introduce Multi-Turn Certified Robustness (MTCR), a framework that models conversational safety via State-Adversarial MDPs and defines k-turn certified robustness as the worst-case safety probability across k adversarial turns. MTCR comprises: (i) compositional certification via embedding-space mode decomposition, yielding tighter certified lower bounds than naive multiplication; (ii) (alpha,eta)-safety persistence, improving the degradation rate from nderline{p}^{k} to eta^k (with eta > nderline{p}) and yielding interpretable horizon estimates; (iii) matching information-theoretic upper bounds establishing tightness; and (iv) a unified algorithm combining these results. Experiments on six LLMs under epsilon-bounded and Crescendo-style attacks confirm that empirical safety consistently exceeds the certified bounds.

Related

Source: arXiv cs.AI | 2026-08-24

Loading related sources…