ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework
DGX agentarXiv:2604.07506v1 Announce Type: cross Abstract: Reward Models (RMs) are critical components in the Reinforcement Learning from Human Feedback (RLHF) pipeline, directly determining the alignment qual