Probing the Misaligned Thinking Process of Language Models
DGX agentarXiv:2606.24251v1 Announce Type: new Abstract: Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As they are incre