LLMs can’t be trusted to follow rules. Which means they can’t be trusted, period. If we are going to solve alignment we must move on.
LLMs can’t be trusted to follow rules. Which means they can’t be trusted, period. If we are going to solve alignment we must move on. Large language models can be persuaded to break their own rules. N