Safety

We’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what user…

We’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for

DGX agentx-post
safetyopenai--x

We’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for measuring how strongly such beliefs shape behavior. https://alignment.openai.com/measuring-reward-seeking/

Related

Source: OpenAI (X) | 2026-07-21

Loading related sources…