ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
arXiv:2505.19241v2 Announce Type: replace-cross Abstract: The recent success in using human preferences to align large language models (LLMs) has significantly improved their performance in various do