MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
DGX agentarXiv:2602.17658v2 Announce Type: replace-cross Abstract: Reward modeling is central to alignment pipelines such as RLHF, RLAIF, and PPO-based policy optimization, yet its reliability is constrained b