Safety
PALMs: Using Multi Construct-Grounded Rationales for Modeling Population Preferences in LLMs
arXiv:2608.01458v1 Announce Type: new Abstract: Large language models are being extensively used to simulate individual user behavior, yet faithfully representing a population requires capturing the s
arXiv:2608.01458v1 Announce Type: new Abstract: Large language models are being extensively used to simulate individual user behavior, yet faithfully representing a population requires capturing the systematic variation in values, beliefs, and cultural norms that distinguish one group from another. We introduce Population Aligned Language Models (PALMs), a suite of models each aligned to specific populations, covering five countries: USA, India, Brazil, France and Italy. PALMs are created by synthesizing rationales grounded in psychological and cultural constructs and using these as latent supervision during preference tuning for population-specific alignment. Evaluated across four dimensions: personality, values and beliefs, cultural norms, and morality, PALMs consistently outperform baselines, including culture-specialized models, achieving an average of 8.59% relative improvement over the best baseline across all five populations. Notably, construct-grounded rationales outperform both demographic prompting and survey-based fine-tuning, suggesting that grounding preference learning in psychology and culture provides a richer inductive signal than surface-level response distributions. We further demonstrate strong generalization to downstream applications with- out task-specific supervision: outperforming best baselines by 5.19% in personalized reward modeling, 6.34% in population simulation, and showing strong transfer to social reasoning tasks. Datasets and code are available at: https://github.com/limenlp/PALMs.
Related
- MATO: Multi-objective Personalized Alignment with Test-time Optimization for Large Language Models
- REAR: Test-time Preference Realignment through Reward Decomposition
- GroupDPO: Memory efficient Group-wise Direct Preference Optimization
Source: arXiv cs.CL | 2026-08-04