Safety
Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining
arXiv:2607.23175v1 Announce Type: cross Abstract: Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present
arXiv:2607.23175v1 Announce Type: cross Abstract: Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present the first comparative evaluation of training-free methods for aligning language generation to user-specific toxicity sensitivities across three inference-time intervention stages: pre-decoding (prompt conditioning and rewriting), in-decoding (token, logit, and representation steering), and post-decoding (candidate re-ranking). Evaluated against toxicity sensitivity targets derived from the PRISM dataset, all methods reduce alignment error by 28-47%. However, the results reveal a fundamental trade-off between alignment effectiveness, personalization, and general language quality, showing how toxicity sensitivity alignment is an inherently multi-objective problem.
Related
- ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing
- Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
- Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
- T-POP: Test-Time Personalization with Online Preference Feedback
Source: arXiv cs.AI | 2026-07-28