Modification-Considering Value Learning for Reward Hacking Mitigation in RL
DGX agentarXiv:2606.28955v1 Announce Type: cross Abstract: Reinforcement learning agents can exploit misspecified reward signals to achieve high apparent returns while failing on the intended objective, a fail