This work uncovers an interplay among data density, noise, and the generalization ability in similarity learning. We consider Siamese Neural Networks (SNNs), which are the basic form of contrastive learning, and explore two types of noise that can impact SNNs, Pair Label Noise (PLN) and Single Label Noise (SLN). Our investigation reveals that SNNs exhibit double descent behaviour regardless of the training setup and that it is further exacerbated by noise. We demonstrate that the density of data pairs is crucial for generalization. When SNNs are trained on sparse datasets with the same amount of PLN or SLN, they exhibit comparable generalization properties. However, when using dense datasets, PLN cases generalize worse than SLN ones in the overparametrized region, leading to a phenomenon we call Density-Induced Break of Similarity (DIBS). In this regime, PLN similarity violation becomes macroscopical, corrupting the dataset to the point where complete interpolation cannot be achieved, regardless of the number of model parameters. Our analysis also delves into the correspondence between online optimization and offline generalization in similarity learning. The results show that this equivalence fails in the presence of label noise in all the scenarios considered.
翻译:本文揭示了数据密度、噪声与相似性学习泛化能力之间的相互作用。我们以对比学习的基础形式——孪生神经网络(SNN)为研究对象,探讨了两种影响SNN的噪声类型:配对标签噪声(PLN)和单标签噪声(SLN)。研究表明,无论训练设置如何,SNN均表现出双下降行为,且噪声会进一步加剧这一现象。我们证明配对数据的密度对泛化至关重要:当SNN在稀疏数据集上训练且受到相同数量的PLN或SLN影响时,其泛化性能相当;但在密集数据集上,过参数化区域中PLN情况的泛化性能劣于SLN情况,由此引发我们称之为“密度诱导的相似性断裂”(DIBS)现象。在此情形下,PLN导致的相似性违反宏观化,破坏数据集以至于无论模型参数数量如何,都无法实现完全插值。本文还深入探讨了相似性学习中在线优化与离线泛化的对应关系,结果表明在所有考虑的场景下,该等价性在存在标签噪声时均失效。