From Rebound to Remedy: Understanding and Mitigating Reward Hacking via Representation Engineering
DGX agentarXiv:2604.01476v2 Announce Type: replace-cross Abstract: Reinforcement learning for LLMs is vulnerable to reward hacking, where models exploit shortcuts to maximize reward without solving the intende