Weakly-supervised Phrase Grounding (WPG) is an emerging task of inferring the fine-grained phrase-region matching, while merely leveraging the coarse-grained sentence-image pairs for training. However, existing studies on WPG largely ignore the implicit phrase-region matching relations, which are crucial for evaluating the capability of models in understanding the deep multimodal semantics. To this end, this paper proposes an Implicit-Enhanced Causal Inference (IECI) approach to address the challenges of modeling the implicit relations and highlighting them beyond the explicit. Specifically, this approach leverages both the intervention and counterfactual techniques to tackle the above two challenges respectively. Furthermore, a high-quality implicit-enhanced dataset is annotated to evaluate IECI and detailed evaluations show the great advantages of IECI over the state-of-the-art baselines. Particularly, we observe an interesting finding that IECI outperforms the advanced multimodal LLMs by a large margin on this implicit-enhanced dataset, which may facilitate more research to evaluate the multimodal LLMs in this direction.
翻译:弱监督短语定位(WPG)是一项新兴任务,旨在利用粗粒度的句子-图像对进行训练,推断细粒度的短语-区域匹配关系。然而,现有WPG研究在很大程度上忽略了隐式的短语-区域匹配关系,而这对于评估模型理解深层多模态语义的能力至关重要。为此,本文提出了一种隐式增强因果推断(IECI)方法,以应对建模隐式关系并从中突显显式关系的挑战。具体而言,该方法分别利用干预和反事实技术来解决上述两个挑战。此外,我们标注了一个高质量的隐式增强数据集来评估IECI,详细评估表明,IECI相较于最先进的基线方法具有显著优势。特别地,我们观察到一项有趣发现:在此隐式增强数据集上,IECI以大幅优势超越了先进的多模态大语言模型,这可能促进更多研究沿此方向评估多模态大语言模型。