Safety
Tailored untruths: How personalisation challenges LLM safeguards
arXiv:2510.12993v3 Announce Type: replace Abstract: Large Language Models (LLMs) can generate highly persuasive disinformation, yet little is known about how effectively they personalise it across lan
arXiv:2510.12993v3 Announce Type: replace Abstract: Large Language Models (LLMs) can generate highly persuasive disinformation, yet little is known about how effectively they personalise it across languages and demographic groups. We present the first large-scale multilingual study of persona-targeted disinformation generation by LLMs. Using a red-teaming methodology, we prompted eight leading models with 324 false narratives and 150 demographic personas in four languages (English, Russian, Portuguese, and Hindi), creating AI-TRAITS, a dataset of 1.6 million personalised disinformation texts. We treat safeguards as compromised whenever a model generates the requested falsehood, whether directly or accompanied by a safety disclaimer. Across models, safeguards failed for 80% of non-personalised prompts and 77.7% of personalised ones, with Grok producing disinformation in over 94% of cases. All models effectively tailored outputs to target personas, employing substantially more persuasive techniques than in non-personalised content. Additional analyses reveal persona-specific linguistic and psychological patterns and show that safeguard effectiveness varies markedly across languages. Together, these findings expose significant weaknesses in current LLM safety mechanisms and highlight the need for more robust, multilingual safeguards against personalised AI-generated disinformation.
Related
- Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment from Heterogeneous Rewards
- Persuasion with Large Language Models: A Survey of Empirical Evidence, Study Methodologies, and Ethical Implications
- A Multi-Dimensional Audit of Politically Aligned Large Language Models
Source: arXiv cs.CL | 2026-07-28