Safety
STAIF: A Stage-wise Optimization for Complex Instruction Following
arXiv:2607.22649v1 Announce Type: new Abstract: Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existing alignment m
arXiv:2607.22649v1 Announce Type: new Abstract: Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existing alignment methods, such as DPO, optimize holistic reward signals that often underemphasize strict satisfaction of individual constraints, particularly under out-of-distribution or multi-constraint settings. In this paper, we propose STAIF, a stage-wise optimization framework that decouples the alignment of subjective (soft) constraints from the optimization of objectively verifiable (hard) constraints. Stage 1 applies preference optimization with multiple negative samples to sharpen sensitivity to soft constraints, while Stage 2 applies Reinforcement Learning with Verifiable Rewards (RLVR) to enforce strict compliance with hard constraints. To support this method, we construct STAINSTRUCT, a high-quality bilingual (English, Chinese) dataset of approximately 31,000 complex multi-constraint instructions. Extensive analyses validate the design of STAIF and show state-of-the-art performance on representative benchmarks against strong baselines, as well as genuine generalization.
Related
- Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment
- VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models
- Configurable Reward Model for Balanced Safety Alignment
- ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
Source: arXiv cs.AI | 2026-07-28