Negative sampling stands as a pivotal technique in dense retrieval, essential for training effective retrieval models and significantly impacting retrieval performance. While existing negative sampling methods have made commendable progress by leveraging hard negatives, a comprehensive guiding principle for constructing negative candidates and designing negative sampling distributions is still lacking. To bridge this gap, we embark on a theoretical analysis of negative sampling in dense retrieval. This exploration culminates in the unveiling of the quasi-triangular principle, a novel framework that elucidates the triangular-like interplay between query, positive document, and negative document. Fueled by this guiding principle, we introduce TriSampler, a straightforward yet highly effective negative sampling method. The keypoint of TriSampler lies in its ability to selectively sample more informative negatives within a prescribed constrained region. Experimental evaluation show that TriSampler consistently attains superior retrieval performance across a diverse of representative retrieval models.
翻译:负采样是稠密检索中的关键技术,对训练高效检索模型至关重要,并显著影响检索性能。尽管现有负采样方法通过利用困难负样本取得了显著进展,但在构建负候选集与设计负采样分布方面仍缺乏系统的指导原则。为弥补这一不足,我们对稠密检索中的负采样进行了理论分析,最终揭示了准三角原则——一种阐明查询、正文档与负文档间三角类交互关系的新框架。基于该指导原则,我们提出了TriSampler——一种简洁而高效的负采样方法。其核心在于能够在预设约束区域内选择性采样信息量更大的负样本。实验评估表明,TriSampler在多种代表性检索模型上持续取得了卓越的检索性能。