Safety
Progressing beyond Art Masterpieces or Touristic Cliches: how to assess your LLMs for cultural alignment?
arXiv:2604.25654v1 Announce Type: new Abstract: Although the cultural (mis)alignment of Large Language Models (LLMs) has attracted increasing attention -- often framed in terms of cultural bias -- unt
arXiv:2604.25654v1 Announce Type: new Abstract: Although the cultural (mis)alignment of Large Language Models (LLMs) has attracted increasing attention -- often framed in terms of cultural bias -- until recently there has been limited work on the design and development of datasets for cultural assessment. Here, we review existing approaches to such datasets and identify their main limitations. To address these issues, we propose design guidelines for annotators and report on the construction of a dataset built according to these principles. We further present a series of contrastive experiments conducted with this dataset. The results demonstrate that our design yields test sets with greater discriminative power, effectively distinguishing between models specialized for a given culture and those that are not, ceteris paribus.
Related
- C-Mining: Unsupervised Discovery of Seeds for Cultural Data Synthesis via Geometric Misalignment
- Self-Debias: Self-correcting for Debiasing Large Language Models
- Meet Dynamic Individual Preferences: Resolving Conflicting Human Value with Paired Fine-Tuning
- Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs
- Whose Story Gets Told? Positionality and Bias in LLM Summaries of Life Narratives
Source: arXiv cs.CL | 2026-04-29