Applications
Stop Listening to Me! How Multi-turn Conversations Can Degrade LLM Diagnostic Reasoning
arXiv:2603.11394v2 Announce Type: replace Abstract: Patients and clinicians are increasingly using chatbots powered by large language models (LLMs) for healthcare inquiries. While state-of-the-art LLM
arXiv:2603.11394v2 Announce Type: replace Abstract: Patients and clinicians are increasingly using chatbots powered by large language models (LLMs) for healthcare inquiries. While state-of-the-art LLMs exhibit high performance on static diagnostic reasoning benchmarks, their efficacy across multi-turn conversations, which better reflect real-world usage, has been understudied. In this paper, we evaluate 17 LLMs across three clinical datasets to investigate how partitioning the decision-space into multiple simpler turns of conversation influences their diagnostic reasoning. Specifically, we develop a "stick-or-switch" evaluation framework to measure model conviction (i.e., defending a correct diagnosis or safe abstention against incorrect suggestions) and flexibility (i.e., recognizing a correct suggestion when it is introduced) across conversations. Our experiments reveal the conversation tax, where multi-turn interactions consistently degrade performance when compared to single-shot baselines. Notably, models frequently abandon initial correct diagnoses and safe abstentions to align with incorrect user suggestions. Additionally, several models exhibit blind switching, failing to distinguish between signal and incorrect suggestions.
Related
- AI generates well-liked but templatic empathic responses
- DQA: Diagnostic Question Answering for IT Support
- DYCP: Dynamic Context Pruning for Long-Form Dialogue with LLMs
- Data Selection for Multi-turn Dialogue Instruction Tuning
- Stay Focused: Problem Drift in Multi-Agent Debate
Source: arXiv cs.CL | 2026-04-10