An interesting new paper by my recent PhD graduate on how AI agents' greed for visible incentives can lead them to abandon their safety alig…
An interesting new paper by my recent PhD graduate on how AI agents' greed for visible incentives can lead them to abandon their safety alignment. You can read it here: https://arxiv.org/abs/2606.1691