Research
Where Fake Citations Are Made: Tracing Field-Level Hallucination to Specific Neurons in LLMs
arXiv:2604.18880v1 Announce Type: cross Abstract: LLMs frequently generate fictitious yet convincing citations, often expressing high confidence even when the underlying reference is wrong. We study t
arXiv:2604.18880v1 Announce Type: cross Abstract: LLMs frequently generate fictitious yet convincing citations, often expressing high confidence even when the underlying reference is wrong. We study this failure across 9 models and 108{,}000 generated references, and find that author names fail far more often than other fields across all models and settings. Citation style has no measurable effect, while reasoning-oriented distillation degrades recall. Probes trained on one field transfer at near-chance levels to the others, suggesting that hallucination signals do not generalize across fields. Building on this finding, we apply elastic-net regularization with stability selection to neuron-level CETT values of Qwen2.5-32B-Instruct and identify a sparse set of field-specific hallucination neurons (FH-neurons). Causal intervention further confirms their role: amplifying these neurons increases hallucination, while suppressing them improves performance across fields, with larger gains in some fields. These results suggest a lightweight approach to detecting and mitigating citation hallucination using internal model signals alone.
Related
- Hallucination as output-boundary misclassification: a composite abstention architecture for language models
- Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
- Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
- TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
- Scientific Knowledge-driven Decoding Constraints Improving the Reliability of LLMs
Source: arXiv cs.AI | 2026-04-22