Vision-language pretrained models have seen remarkable success, but their application to safety-critical settings is limited by their lack of interpretability. To improve the interpretability of vision-language models such as CLIP, we propose a multi-modal information bottleneck (M2IB) approach that learns latent representations that compress irrelevant information while preserving relevant visual and textual features. We demonstrate how M2IB can be applied to attribution analysis of vision-language pretrained models, increasing attribution accuracy and improving the interpretability of such models when applied to safety-critical domains such as healthcare. Crucially, unlike commonly used unimodal attribution methods, M2IB does not require ground truth labels, making it possible to audit representations of vision-language pretrained models when multiple modalities but no ground-truth data is available. Using CLIP as an example, we demonstrate the effectiveness of M2IB attribution and show that it outperforms gradient-based, perturbation-based, and attention-based attribution methods both qualitatively and quantitatively.
翻译:视觉-语言预训练模型已取得显著成功,但其在安全关键场景中的应用受限于可解释性不足。为提升CLIP等视觉-语言模型的可解释性,我们提出多模态信息瓶颈(M2IB)方法,该方法通过学习压缩无关信息并保留相关视觉与文本特征的潜在表征。我们展示了如何将M2IB应用于视觉-语言预训练模型的归因分析,提升归因准确性,并改善此类模型在医疗等安全关键领域的可解释性。关键区别在于,与常用的单模态归因方法不同,M2IB无需依赖真实标签,从而能在多模态可用但无真实数据的情况下审计视觉-语言预训练模型的表征。以CLIP为例,我们验证了M2IB归因的有效性,并证明其在定性与定量层面均优于基于梯度、扰动和注意力机制的归因方法。