Despite the success of multimodal learning in cross-modal retrieval task, the remarkable progress relies on the correct correspondence among multimedia data. However, collecting such ideal data is expensive and time-consuming. In practice, most widely used datasets are harvested from the Internet and inevitably contain mismatched pairs. Training on such noisy correspondence datasets causes performance degradation because the cross-modal retrieval methods can wrongly enforce the mismatched data to be similar. To tackle this problem, we propose a Meta Similarity Correction Network (MSCN) to provide reliable similarity scores. We view a binary classification task as the meta-process that encourages the MSCN to learn discrimination from positive and negative meta-data. To further alleviate the influence of noise, we design an effective data purification strategy using meta-data as prior knowledge to remove the noisy samples. Extensive experiments are conducted to demonstrate the strengths of our method in both synthetic and real-world noises, including Flickr30K, MS-COCO, and Conceptual Captions.
翻译:尽管多模态学习在跨模态检索任务中取得了成功,但其显著进展依赖于多媒体数据之间的正确对应关系。然而,收集此类理想数据既昂贵又耗时。实际上,大多数广泛使用的数据集是从互联网获取的,不可避免地包含不匹配的样本对。在这种带有噪声的对应数据集上进行训练会导致性能下降,因为跨模态检索方法可能错误地强制不匹配的数据具有相似性。为解决这一问题,我们提出了一种元相似度纠正网络(Meta Similarity Correction Network, MSCN),以提供可靠的相似度分数。我们将二元分类任务视为一个元过程,鼓励MSCN从正负元数据中学习判别能力。为了进一步减轻噪声的影响,我们设计了一种有效的数据净化策略,利用元数据作为先验知识去除噪声样本。通过大量实验,我们在合成噪声和真实噪声场景(包括Flickr30K、MS-COCO和Conceptual Captions数据集)中展示了所提方法的优势。