The Representation-Rationalizability Tradeoff in Reward Learning
DGX agentarXiv:2606.00291v1 Announce Type: cross Abstract: In RLHF, each training example contains a prompt x and two candidate responses y,y', and annotators provide pairwise preferences between these respons