Safety
agreed. RL is not (at least by itself) the way to alignment
agreed. RL is not (at least by itself) the way to alignment Yoshua Bengio says Reinforcement Learning is a dangerous path for building superintelligence It can create systems with hidden goals, reward
agreed. RL is not (at least by itself) the way to alignment Yoshua Bengio says Reinforcement Learning is a dangerous path for building superintelligence It can create systems with hidden goals, reward hacking, and behavior that goes against what humans actually want "an AI that doesn't care about outcomes can't be corrupted by them"
Source: Gary Marcus (X) | 2026-05-13