Research
Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
arXiv:2502.07963v4 Announce Type: replace-cross Abstract: Medical research faces well-documented challenges in translating novel treatments into clinical practice. Publishing incentives encourage rese
arXiv:2502.07963v4 Announce Type: replace-cross Abstract: Medical research faces well-documented challenges in translating novel treatments into clinical practice. Publishing incentives encourage researchers to present "positive" findings, even when empirical results are equivocal. Consequently, it is well-documented that authors often spin study results, especially in article abstracts. Such spin can influence clinician interpretation of evidence and may affect patient care decisions. In this study, we ask whether the interpretation of trial results offered by Large Language Models (LLMs) is similarly affected by spin. This is important since LLMs are increasingly being used to trawl through and synthesize published medical evidence. We evaluated 22 LLMs and found that they are across the board more susceptible to spin than humans. They might also propagate spin into their outputs: We find evidence, e.g., that LLMs implicitly incorporate spin into plain language summaries that they generate. We also find, however, that LLMs are generally capable of recognizing spin, and can be prompted in a way to mitigate spin's impact on LLM outputs.
Related
- Extracting Breast Cancer Phenotypes from Clinical Notes: Comparing LLMs with Classical Ontology Methods
- LLMs Struggle with Abstract Meaning Comprehension More Than Expected
- Evaluating LLMs as Human Surrogates in Controlled Experiments
- Overstating Attitudes, Ignoring Networks: LLM Biases in Simulating Misinformation Susceptibility
Source: arXiv cs.AI | 2026-04-23