Multimodal entity linking (MEL) task, which aims at resolving ambiguous mentions to a multimodal knowledge graph, has attracted wide attention in recent years. Though large efforts have been made to explore the complementary effect among multiple modalities, however, they may fail to fully absorb the comprehensive expression of abbreviated textual context and implicit visual indication. Even worse, the inevitable noisy data may cause inconsistency of different modalities during the learning process, which severely degenerates the performance. To address the above issues, in this paper, we propose a novel Multi-GraIned Multimodal InteraCtion Network $\textbf{(MIMIC)}$ framework for solving the MEL task. Specifically, the unified inputs of mentions and entities are first encoded by textual/visual encoders separately, to extract global descriptive features and local detailed features. Then, to derive the similarity matching score for each mention-entity pair, we device three interaction units to comprehensively explore the intra-modal interaction and inter-modal fusion among features of entities and mentions. In particular, three modules, namely the Text-based Global-Local interaction Unit (TGLU), Vision-based DuaL interaction Unit (VDLU) and Cross-Modal Fusion-based interaction Unit (CMFU) are designed to capture and integrate the fine-grained representation lying in abbreviated text and implicit visual cues. Afterwards, we introduce a unit-consistency objective function via contrastive learning to avoid inconsistency and model degradation. Experimental results on three public benchmark datasets demonstrate that our solution outperforms various state-of-the-art baselines, and ablation studies verify the effectiveness of designed modules.
翻译:多模态实体链接(MEL)任务旨在将模糊指称解析至多模态知识图谱,近年来受到广泛关注。尽管已有大量研究探索多模态间的互补效应,但现有方法仍难以充分吸收简略文本语境与隐含视觉指示的综合表达能力。更严重的是,学习过程中难以避免的噪声数据会导致不同模态间的不一致性,从而严重降低模型性能。针对上述问题,本文提出名为多粒度多模态交互网络(MIMIC)的新型框架来解决MEL任务。具体而言,首先分别通过文本/视觉编码器对指称与实体的统一输入进行编码,提取全局描述特征与局部细节特征。随后,为获取每个指称-实体对的相似度匹配分数,我们设计了三个交互单元,以全面探索实体与指称特征间的模态内交互与模态间融合。其中,基于文本的全局-局部交互单元(TGLU)、基于视觉的双重交互单元(VDLU)和跨模态融合交互单元(CMFU)三个模块专门设计用于捕获并整合简略文本与隐含视觉线索中的细粒度表示。此外,我们通过对比学习引入单元一致性目标函数以避免不一致性与模型退化。在三个公开基准数据集上的实验结果表明,我们的方案优于多种最先进基线模型,消融研究验证了各设计模块的有效性。