Safety
Disentangling Optimization Scale from Preference Scale in DPO
arXiv:2608.27032v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is a widely used objective for aligning language models from preference data, with the coefficient eta commonly int
arXiv:2608.27032v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is a widely used objective for aligning language models from preference data, with the coefficient eta commonly interpreted as controlling the KL constraint to a reference policy. We show that eta entangles two distinct roles: it governs the effective inverse preference-noise scale and simultaneously rescales the optimization dynamics, coupling this scale with the effective step size. As a consequence, at a fixed learning rate the achieved policy deviation is non-monotone in eta: it vanishes in a dead zone at small eta, reaches a peak at an intermediate value, and decreases again for larger eta. Moreover, standard DPO loss values are not comparable across eta: runs with nearly identical loss curves can differ several-fold in KL divergence from the reference model. This entanglement obscures the role of eta, increases sensitivity to hyperparameter choices, and complicates learning-rate scheduling. We propose a centered-softplus reformulation that is argmin-equivalent to DPO for eta>0, while making the inverse preference-noise-scale and learning-rate effects explicit and independently tunable. The normalized centered-softplus objective also admits a continuous etao0 endpoint that reduces to a linear preference-margin objective.
Related
- Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models
- TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
- GroupDPO: Memory efficient Group-wise Direct Preference Optimization
- TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization
Source: arXiv cs.LG | 2026-08-28