Research
Mitigating Multimodal Hallucination via Phase-wise Self-reward
arXiv:2604.17982v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) still struggle with vision hallucination, where generated responses are inconsistent with the visual input. Exist
arXiv:2604.17982v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) still struggle with vision hallucination, where generated responses are inconsistent with the visual input. Existing methods either rely on large-scale annotated data for fine-tuning, which incurs massive computational overhead, or employ static post-hoc strategies that overlook the dynamic nature of hallucination emergence. To address these, we introduce a new self-rewarding framework, enabling dynamic hallucination mitigation at inference time without external supervision. On the empirical side, we reveal that visual hallucination exhibits phase-wise dynamic patterns, peaking at the onset of each semantic phase. Drawing on these insights, we propose extbf{PSRD} (extbf{Phase-wise extbf{S}elf-extbf{R}eward extbf{D}ecoding) for online hallucination correction guided by phase-wise self-reward signals. To reduce the cost of repeated self-evaluation during decoding, we distill the hallucination guidance signal from LVLMs into a lightweight reward model. The reward model subsequently provides on-the-fly guidance for targeted intervention during the decoding process, enabling precise hallucination suppression. The proposed PSRD significantly reduces the hallucination rate of LLaVA-1.5-7B by 50.0% and consistently outperforms existing post-hoc methods across five hallucination evaluation benchmarks for four LVLMs. Further analysis confirms that PSRD effectively mitigates hallucination propagation and achieves a highly controllable trade-off between strong performance and inference efficiency.
Related
- Countering the Over-Reliance Trap: Mitigating Object Hallucination for LVLMs via a Self-Validation Framework
- HTDC: Hesitation-Triggered Differential Calibration for Mitigating Hallucination in Large Vision-Language Models
- Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding
- Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
- VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning
- TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection
Source: arXiv cs.CL | 2026-04-21