We’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what user…
DGX agentWe’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for