MACD: Model-Aware Contrastive Decoding via Counterfactual Data
DGX agentarXiv:2602.01740v3 Announce Type: replace Abstract: Video language models (Video-LLMs) are prone to hallucinations, generating plausible but ungrounded content when visual evidence is weak, ambiguous,