Research
HTDC: Hesitation-Triggered Differential Calibration for Mitigating Hallucination in Large Vision-Language Models
arXiv:2604.12115v1 Announce Type: new Abstract: Large vision-language models (LVLMs) achieve strong multimodal performance, but still suffer from hallucinations caused by unstable visual grounding and
arXiv:2604.12115v1 Announce Type: new Abstract: Large vision-language models (LVLMs) achieve strong multimodal performance, but still suffer from hallucinations caused by unstable visual grounding and over-reliance on language priors. Existing training-free decoding methods typically apply calibration at every decoding step, introducing unnecessary computation and potentially disrupting stable predictions. We address this problem by identifying layer-wise hesitation, a simple signal of grounding instability reflected by fluctuations in token preference across intermediate layers. Based on this observation, we propose Hesitation-Triggered Differential Calibration (HTDC), a training-free decoding framework that preserves standard full-branch inference and activates calibration only at hesitation-prone steps. When triggered, HTDC contrasts the full branch with two lightweight probes, a visual-nullification probe and a semantic-nullification probe, to suppress hallucination-prone candidates while avoiding unnecessary intervention on stable steps. Experiments on representative hallucination benchmarks show that HTDC consistently reduces hallucinations while maintaining strong task accuracy, achieving a favorable trade-off between effectiveness and computational overhead.
Related
- Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
- VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning
- Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding
- TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection
- HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
- Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
Source: arXiv cs.CV | 2026-04-15