Deep Metric Learning (DML) methods aim at learning an embedding space in which distances are closely related to the inherent semantic similarity of the inputs. Previous studies have shown that popular benchmark datasets often contain numerous wrong labels, and DML methods are susceptible to them. Intending to study the effect of realistic noise, we create an ontology of the classes in a dataset and use it to simulate semantically coherent labeling mistakes. To train robust DML models, we propose ProcSim, a simple framework that assigns a confidence score to each sample using the normalized distance to its class representative. The experimental results show that the proposed method achieves state-of-the-art performance on the DML benchmark datasets injected with uniform and the proposed semantically coherent noise.
翻译:深度度量学习(DML)方法旨在学习一个嵌入空间,在该空间中距离与输入的内在语义相似性密切相关。以往研究表明,常用基准数据集通常包含大量错误标签,而DML方法易受其影响。为研究真实噪声的影响,我们构建了数据集中类别的本体,并利用其模拟语义一致的标注错误。为训练鲁棒的DML模型,我们提出ProcSim这一简单框架,通过计算样本到其类别代表的归一化距离为每个样本分配置信度分数。实验结果表明,该方法在注入均匀噪声和所提出的语义一致噪声的DML基准数据集上达到了最先进的性能。