Noisy correspondence that refers to mismatches in cross-modal data pairs, is prevalent on human-annotated or web-crawled datasets. Prior approaches to leverage such data mainly consider the application of uni-modal noisy label learning without amending the impact on both cross-modal and intra-modal geometrical structures in multimodal learning. Actually, we find that both structures are effective to discriminate noisy correspondence through structural differences when being well-established. Inspired by this observation, we introduce a Geometrical Structure Consistency (GSC) method to infer the true correspondence. Specifically, GSC ensures the preservation of geometrical structures within and between modalities, allowing for the accurate discrimination of noisy samples based on structural differences. Utilizing these inferred true correspondence labels, GSC refines the learning of geometrical structures by filtering out the noisy samples. Experiments across four cross-modal datasets confirm that GSC effectively identifies noisy samples and significantly outperforms the current leading methods.
翻译:噪声对应指跨模态数据对中的不匹配现象,在人工标注或网络爬取的数据集中普遍存在。现有利用此类数据的方法主要考虑单模态噪声标签学习的应用,而未修正其对多模态学习中跨模态与模态内几何结构的影响。实际上,我们发现当这两种结构被良好建立时,皆能通过结构差异有效判别噪声对应。受此启发,我们提出几何结构一致性方法以推断真实对应关系。具体而言,GSC确保模态内与模态间几何结构的保持,从而基于结构差异实现噪声样本的精准判别。利用这些推断出的真实对应标签,GSC通过滤除噪声样本来优化几何结构的学习。在四个跨模态数据集上的实验证实,GSC能有效识别噪声样本,并显著优于当前主流方法。