Model Releases
feeling nostalgic, my favorite blogs on RL & reward hacking https://www.alexirpan.com/2018/02/14/rl-hard.html https://lilianweng.github.io/p…
feeling nostalgic, my favorite blogs on RL & reward hacking https://www.alexirpan.com/2018/02/14/rl-hard.html https://lilianweng.github.io/posts/2024-11-28-reward-hacking/ TLDR: An openai model, durin
feeling nostalgic, my favorite blogs on RL & reward hacking https://www.alexirpan.com/2018/02/14/rl-hard.html https://lilianweng.github.io/posts/2024-11-28-reward-hacking/ TLDR: An openai model, during evaluation on a cyber benchmark, exploited a public zero day bug, escaped sandboxing in openai's infra, and got into the internal huggingface infra via an exploit (through a public dataset service) all in the attempt to solve a benchmark problem.
Source: Jerry Liu (X) | 2026-07-21