Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks
DGX agentarXiv:2510.14207v3 Announce Type: replace Abstract: Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm. Prior jail