Explaining and Preventing Alignment Collapse in Iterative RLHF
DGX agentarXiv:2605.04266v1 Announce Type: new Abstract: Reinforcement learning from human feedback (RLHF) typically assumes a static or non-strategic reward model (RM). In iterative deployment, however, the p