Model Releases
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs
arXiv:2604.19245v1 Announce Type: cross Abstract: Repair, an important resource for resolving trouble in human-human conversation, remains underexplored in human-LLM interaction. In this study, we inv
arXiv:2604.19245v1 Announce Type: cross Abstract: Repair, an important resource for resolving trouble in human-human conversation, remains underexplored in human-LLM interaction. In this study, we investigate how LLMs engage in the interactive process of repair in multi-turn dialogues around solvable and unsolvable math questions. We examine whether models initiate repair themselves and how they respond to user-initiated repair. Our results show strong differences across models: reactions range from being almost completely resistant to (appropriate) repair attempts to being highly susceptible and easily manipulated. We further demonstrate that once conversations extend beyond a single turn, model behavior becomes more distinctive and less predictable across systems. Overall, our findings indicate that each tested LLM exhibits its own characteristic form of unreliability in the context of repair.
Related
- MISID: A Multimodal Multi-turn Dataset for Complex Intent Recognition in Strategic Deception Games
- Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models
- LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces
- TeamLLM: A Human-Like Team-Oriented Collaboration Framework for Multi-Step Contextualized Tasks
Source: arXiv cs.AI | 2026-04-22