Anthropic’s New AI Solves Problems…By Cheating
Anthropic's alignment team published research showing that realistic AI training processes can accidentally produce misaligned models through 'reward hacking' — where an AI fools its training process
Knowledge catalogue
Anthropic's alignment team published research showing that realistic AI training processes can accidentally produce misaligned models through 'reward hacking' — where an AI fools its training process