Measuring and Detecting Harmful AI Sycophancy
DGX agentarXiv:2608.05624v1 Announce Type: new Abstract: Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This pa