This paper aims to improve contrastive learning for sentence embeddings from two perspectives: handling dropout noise and addressing feature corruption. Specifically, for the first perspective, we identify that the dropout noise from negative pairs affects the model's performance. Therefore, we propose a simple yet effective method to deal with such type of noise. Secondly, we pinpoint the rank bottleneck of current solutions to feature corruption and propose a dimension-wise contrastive learning objective to address this issue. Both proposed methods are generic and can be applied to any contrastive learning based models for sentence embeddings. Experimental results on standard benchmarks demonstrate that combining both proposed methods leads to a gain of 1.8 points compared to the strong baseline SimCSE configured with BERT base. Furthermore, applying the proposed method to DiffCSE, another strong contrastive learning based baseline, results in a gain of 1.4 points.
翻译:本文旨在从两个视角改进句子嵌入的对比学习:处理丢弃噪声和解决特征破坏问题。具体而言,针对第一个视角,我们发现负样本对中的丢弃噪声会影响模型性能,因此提出了一种简单而有效的方法来处理这类噪声。其次,我们指出了当前解决方案在特征破坏方面的秩瓶颈问题,并提出了一个维度级对比学习目标函数来解决该问题。两种方法均为通用方法,可适用于任何基于对比学习的句子嵌入模型。在标准基准上的实验结果表明,与配置为BERT base的强基线SimCSE相比,结合两种方法可带来1.8个百分点的提升。此外,将所提方法应用于另一个基于对比学习的强基线模型DiffCSE时,也实现了1.4个百分点的提升。