Visual evidence selection is a critical component of multimodal retrieval-augmented generation (RAG), yet existing methods typically rely on semantic relevance or surface-level similarity, which are often misaligned with the actual utility of visual evidence for downstream reasoning. We reformulate multimodal evidence selection from an information-theoretic perspective by defining evidence utility as the information gain induced on a model's output distribution. To overcome the intractability of answer-space optimization, we introduce a latent notion of evidence helpfulness and theoretically show that, under mild assumptions, ranking evidence by information gain on this latent variable is equivalent to answer-space utility. We further propose a training-free, surrogate-accelerated framework that efficiently estimates evidence utility using lightweight multimodal models. Experiments on MRAG-Bench and Visual-RAG across multiple model families demonstrate that our method consistently outperforms state-of-the-art RAG baselines while achieving substantial reductions in computational cost.
翻译:视觉证据选择是多模态检索增强生成的关键组成部分,然而现有方法通常依赖于语义相关性或表层相似性,这些往往与视觉证据在下游推理中的实际效用存在偏差。我们从信息论角度重新定义多模态证据选择,将证据效用定义为模型输出分布上的信息增益。为克服答案空间优化的难解性,我们引入证据帮助性的隐变量概念,并从理论上证明在温和假设条件下,基于该隐变量信息增益排序证据等价于答案空间效用。我们进一步提出无需训练、代理加速的框架,利用轻量级多模态模型高效估计证据效用。在MRAG-Bench和Visual-RAG数据集上跨多个模型族的实验表明,我们的方法在显著降低计算成本的同时,持续优于最先进的检索增强生成基线方法。