PsychoPass: Geometric Profiling of Multi-Turn Adversarial LLM Conversations
DGX agentarXiv:2606.03136v1 Announce Type: cross Abstract: Multi-turn jailbreak attacks on large language models (LLMs) reveal a mismatch in current guardrails: they operate on individual turns, while attacks