Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
arXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabil