Model Releases
Asymmetries in Spontaneous and Instructed Deception
arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed decepti
arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct. We compared these two deception settings through direction geometry, cross-setting classifiers, and cross-setting steering. We found the two deception settings share a component of direction (cosine of approximately 0.5) and an asymmetry in the transfer between settings regarding detection and causation. Spontaneous trained classifiers performed better on instructed data than vice versa, and instructed derived directions performed better at steering spontaneous prompts than vice versa. Likewise the best token position to derive steering vectors from differed from the best token position to train and apply classifiers.
Related
- Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives
- Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies
- LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces
Source: arXiv cs.AI | 2026-09-02