Model Releases
Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory
arXiv:2607.23927v1 Announce Type: new Abstract: A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity i
arXiv:2607.23927v1 Announce Type: new Abstract: A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity is called reality monitoring, and its failures are linked to hallucinations, delusions, and confabulation, yet whether LLMs possess it remains untested. Here we show, across two experiments and six LLMs, that source attribution depends on how conversational memory is structured: ceiling accuracy for self-generated content under minimal memory demands reverses to a fragile external-item advantage once episodic delay removes that shortcut. Feedback exposes two failures: in some models, internal and external judgments swap; in others, accuracy improves while confidence decouples from correctness, dissociations invisible to existing benchmarks. Across models, this pattern implicates active, not aggregate, parameter count. This suggests that as AI systems take on autonomous, multi-turn roles, evaluating what they know is not enough: tracking where that knowledge came from may matter equally.
Related
- Self-Awareness before Action: Mitigating Logical Inertia via Proactive Cognitive Awareness
- Process Matters more than Output for Distinguishing Humans from Machines
- MERMAID: Memory-Enhanced Retrieval and Reasoning with Multi-Agent Iterative Knowledge Grounding for Veracity Assessment
- Towards Scalable Lifelong Knowledge Editing with Selective Knowledge Suppression
Source: arXiv cs.AI | 2026-07-28