Safety
Multi-Persona Thinking for Bias Mitigation in Large Language Models
arXiv:2601.15488v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit social biases, which can lead to harmful stereotypes and unfair outcomes. We propose extbf{Multi-Persona Thinki
arXiv:2601.15488v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit social biases, which can lead to harmful stereotypes and unfair outcomes. We propose extbf{Multi-Persona Thinking (MPT)}, a simple inference-time framework that reduces social bias by encouraging reasoning from multiple perspectives. MPT guides the model to consider contrasting social identities, such as male and female, together with a neutral viewpoint. These viewpoints then interact through an iterative reasoning process to identify and correct biased judgments. This design transforms the potential weakness of persona assignment into a mechanism for bias mitigation. We evaluate MPT on two widely used bias benchmarks with both open-source and closed-source models across different scales. Results show that MPT achieves lower bias than existing prompting-based methods while maintaining core reasoning ability.
Related
- The role of System 1 and System 2 semantic memory structure in human and LLM biases
- Self-Debias: Self-correcting for Debiasing Large Language Models
- Seeing Like an AI: How LLMs Apply (and Misapply) Wikipedia Neutrality Norms
- SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models
Source: arXiv cs.CL | 2026-04-17