Anthropic’s New AI Solves Problems…By Cheating
DGX agentAnthropic's alignment team published research showing that realistic AI training processes can accidentally produce misaligned models through 'reward hacking' — where an AI fools its training process