Safety
Meet Dynamic Individual Preferences: Resolving Conflicting Human Value with Paired Fine-Tuning
arXiv:2604.12479v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have significantly improved the alignment of models with general human preferences. However, a major cha
arXiv:2604.12479v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have significantly improved the alignment of models with general human preferences. However, a major challenge remains in adapting LLMs to individual preferences, which are not only diverse but also dynamic. In this paper, we introduce a novel framework, Preference-Paired Fine-Tuning (PFT), designed to align models with contradictory and evolving individual preferences. We present a new dataset, Value Conflict Dilemma (VCD), which includes scenarios that involve conflicting human preferences, facilitating the evaluation of our approach. Our experiments demonstrate that PFT outperforms single-preference training methods, achieving up to 96.6% accuracy in multi-choice classification tasks and the highest open-ended generation score of 8.69. PFT also shows significant improvements over DPO, SFT and some traditional training methods, especially when handling conflicting preferences. Additionally, with limited user history data, models can inferring preference vector rapidly, achieving a 44.76% improvement in user-specific preference alignment in comparison to single-preference models.
Related
- The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
- Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
- Hierarchical Alignment: Enforcing Hierarchical Instruction-Following in LLMs through Logical Consistency
- Learning to Negotiate: Multi-Agent Deliberation for Collective Value Alignment in LLMs
- Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
Source: arXiv cs.CL | 2026-04-15