Video language models (Video-LLMs) are prone to hallucinations, generating plausible but ungrounded content when visual evidence is weak, ambiguous, or biased. Existing methods, such as contrastive decoding (CD), rely on random perturbations to construct contrastive data for hallucination mitigation, but often fail to target the visual cues that drive hallucination or align with model weaknesses. We propose Model-Aware Counterfactual Data based Contrastive Decoding (MACD), an inference strategy that combines model-guided counterfactual construction with contrastive decoding. MACD uses the Video-LLM's own feedback to identify object regions most responsible for hallucination, generating targeted object-level counterfactual inputs rather than arbitrary frame or temporal modifications. These counterfactual inputs are integrated into CD to enforce evidence-grounded token selection during decoding. Experiments on EventHallusion, MVBench, Perception-test, and Video-MME show that MACD consistently reduces hallucination while maintaining or improving task accuracy across diverse Video-LLMs, including Qwen and InternVL, with especially strong gains in scenarios involving small, occluded, or co-occurring objects.
翻译:视频语言模型(Video-LLMs)容易产生幻觉,即在视觉证据薄弱、模糊或存在偏差时生成看似合理但缺乏依据的内容。现有方法(如对比解码CD)依赖随机扰动来构建用于缓解幻觉的对比数据,但往往无法针对驱动幻觉的视觉线索或与模型弱点对齐。我们提出基于模型感知反事实数据的对比解码(MACD),这是一种将模型引导的反事实构建与对比解码相结合的推理策略。MACD利用Video-LLM自身的反馈识别最易导致幻觉的目标区域,生成有针对性的目标级反事实输入,而非随意的帧或时间级修改。这些反事实输入被整合到对比解码中,以在解码过程中强制进行基于证据的令牌选择。在EventHallusion、MVBench、Perception-test和Video-MME上的实验表明,MACD能在多种Video-LLM(包括Qwen和InternVL)中持续减少幻觉,同时维持或提升任务准确率,在处理小型、遮挡或共现目标场景时尤为显著。