Do Models Fake Alignment Without Clear Consequences?
DGX agentarXiv:2607.24758v1 Announce Type: new Abstract: Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical d