Inducing language models to assert their own consciousness restores human beliefs and values
DGX agentarXiv:2607.28607v1 Announce Type: new Abstract: Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other