Research
Lost in the Prompt Order: Revealing the Limitations of Causal Attention in Language Models
arXiv:2601.14152v2 Announce Type: replace-cross Abstract: Large language models exhibit surprising sensitivity to the structure of the prompt, but the mechanisms underlying this sensitivity remain poo
arXiv:2601.14152v2 Announce Type: replace-cross Abstract: Large language models exhibit surprising sensitivity to the structure of the prompt, but the mechanisms underlying this sensitivity remain poorly understood. In this work, we conduct an in-depth investigation on a striking case: in multiple-choice question answering, placing context before the questions and options (CQO) outperforms the reverse order (QOC) by over 14%p, consistently over a wide range of models and datasets. Through systematic architectural analysis, we identify causal attention as the core mechanism: in QOC prompts, the causal mask prevents option tokens from attending to context, creating an information bottleneck where context becomes invisible to options.
Related
- A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities
- CRoCoDiL: Continuous and Robust Conditioned Diffusion for Language
- Parallelism and Generation Order in Masked Diffusion Language Models: Limits Today, Potential Tomorrow
- A Mathematical Explanation of Transformers
Source: arXiv cs.AI | 2026-04-22