Safety

LLMs can’t be trusted to follow rules. Which means they can’t be trusted, period. If we are going to solve alignment we must move on.

LLMs can’t be trusted to follow rules. Which means they can’t be trusted, period. If we are going to solve alignment we must move on. Large language models can be persuaded to break their own rules. N

DGX agentx-post
safetygary-marcus--x

LLMs can’t be trusted to follow rules. Which means they can’t be trusted, period. If we are going to solve alignment we must move on. Large language models can be persuaded to break their own rules. Not with fancy code. With actual persuasion. The authors tested classic persuasion principles, such as authority, commitment, liking, reciprocity, scarcity, social proof, and unity, analysing over 126,000 conversati…

Source: Gary Marcus (X) | 2026-06-22

Loading related sources…