Model Releases
Mimicry without understanding: the origins of decision bias in large language models
arXiv:2608.12339v1 Announce Type: cross Abstract: Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms through whi
arXiv:2608.12339v1 Announce Type: cross Abstract: Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms through which such biases can be generated even when human preferences (in the training data) are not biased or when they are correctly categorized as being biased. The first is faulty mimicry of preferences based on human behavior: this involves LLMs inferring human preferences even when behaviors are logically unrelated to preferences. The second is mimicry of explicitly biased human behaviors. In four studies focusing on economic biases, we find that ChatGPT-4o and Qwen exhibited social proof biases even when prompted with reports of human behaviors that were clearly non-indicative of individuals' actual preferences. LLMs also displayed loss aversion when it was explicitly described as a bias. Indeed, when prompted with detailed scientific reports, the extent of the bias (i.e., loss aversion) in the scientific report predicted LLMs' own subsequent bias. Scientific papers of biases can thus become self-fulfilling prophecies, at least when it comes to LLMs' responses. The current study goes beyond fleshing out LLM biases and sheds light on the underlying component processes.
Related
- Digital Skin, Digital Bias: Uncovering Tone-Based Biases in LLMs and Emoji Embeddings
- Invisible Influences: Investigating Implicit Intersectional Biases through Persona Engineering in Large Language Models
- Insidious by Design: Implications of Large Language Model algorithmic bias for the Global South
- Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents
Source: arXiv cs.AI | 2026-08-14