Safety
The Imperfective Paradox in Large Language Models
arXiv:2601.09373v2 Announce Type: replace Abstract: Do Large Language Models (LLMs) genuinely grasp the compositional semantics of events, or do they rely on surface-level probabilistic heuristics? We
arXiv:2601.09373v2 Announce Type: replace Abstract: Do Large Language Models (LLMs) genuinely grasp the compositional semantics of events, or do they rely on surface-level probabilistic heuristics? We investigate the Imperfective Paradox, a logical phenomenon where the past progressive aspect entails event realization for activities (e.g., running o ran) but not for accomplishments (e.g., building nrightarrow built). We introduce ImperfectiveNLI, a diagnostic dataset designed to probe this distinction across diverse semantic classes. Evaluating state-of-the-art open-weight models, we uncover a pervasive Teleological Bias: models systematically hallucinate completion for goal-oriented events, even overriding explicit textual cancellation. Prompting interventions partially reduce this bias but trigger a calibration crisis, causing models to incorrectly reject valid entailments for atelic verbs. Representational analyses further show that while internal embeddings often distinguish progressive from simple past forms, inference decisions are dominated by strong priors about goal attainment. Taken together, our findings indicate that these current open-weight LLMs operate as predictive narrative engines rather than faithful logical reasoners, and that resolving aspectual inference requires moving beyond prompting toward structurally grounded alignment.
Related
- Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement
- How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects
- Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
- DiffCoT: Diffusion-styled Chain-of-Thought Reasoning in LLMs
Source: arXiv cs.CL | 2026-04-23