Research
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
arXiv:2505.09591v3 Announce Type: replace-cross Abstract: Despite remarkable progress in recent years, Vision Language Models (VLMs) remain prone to overconfidence and hallucinations on tasks such as
arXiv:2505.09591v3 Announce Type: replace-cross Abstract: Despite remarkable progress in recent years, Vision Language Models (VLMs) remain prone to overconfidence and hallucinations on tasks such as Visual Question Answering (VQA) and Visual Reasoning. Bayesian methods can potentially improve reliability by helping models predict selectively, that is, models respond only when they are sufficiently confident. Unfortunately, such approaches can be costly and ineffective for large models, and there exists little evidence to show otherwise for multimodal applications. Here, we show for the first time the effectiveness and competitive edge of variational Bayes for selective prediction in VQA. We build on recent advances in variational methods for deep learning and propose an extension called "Variational VQA". This method improves calibration and yields significant gains for selective prediction on VQA and Visual Reasoning, particularly when the error tolerance is low (leq 1%). Often, just one posterior sample yields more reliable answers than those given by models trained with AdamW. In addition, we propose a new risk-averse selector that outperforms standard sample averaging by considering the variance of predictions. Overall, we present compelling evidence that variational learning is a viable option to make large VLMs safer and more trustworthy.
Related
- Distorted or Fabricated? A Survey on Hallucination in Video LLMs
- Countering the Over-Reliance Trap: Mitigating Object Hallucination for LVLMs via a Self-Validation Framework
- TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
- Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
- Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
- VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning
Source: arXiv cs.AI | 2026-04-14