Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
DGX agentarXiv:2604.27093v1 Announce Type: cross Abstract: Current LLM safety alignment techniques improve model robustness against adversarial attacks, but overlook whether and how LLMs can recover helpfulnes