Model Releases
DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories
arXiv:2604.20443v1 Announce Type: cross Abstract: Large Language Models (LLMs) have been shown to possess Theory of Mind (ToM) abilities. However, it remains unclear whether this stems from robust rea
arXiv:2604.20443v1 Announce Type: cross Abstract: Large Language Models (LLMs) have been shown to possess Theory of Mind (ToM) abilities. However, it remains unclear whether this stems from robust reasoning or spurious correlations. We introduce DialToM, a human-verified benchmark built from natural human dialogue using a multiple-choice framework. We evaluate not only mental state prediction (Literal ToM) but also the functional utility of these states (Functional ToM) through Prospective Diagnostic Forecasting -- probing whether models can identify state-consistent dialogue trajectories solely from mental-state profiles. Our results reveal a significant reasoning asymmetry: while LLMs excel at identifying mental states, most (except for Gemini 3 Pro) fail to leverage this understanding to forecast social trajectories. Additionally, we find only weak semantic similarities between human and LLM-generated inferences. To facilitate reproducibility, the DialToM dataset and evaluation code are publicly available at https://github.com/Stealth-py/DialToM.
Related
- Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
- METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models
- Robust Reasoning Benchmark
- Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
Source: arXiv cs.AI | 2026-04-23