In this paper, we introduce CuSINeS, a negative sampling approach to enhance the performance of Statutory Article Retrieval (SAR). CuSINeS offers three key contributions. Firstly, it employs a curriculum-based negative sampling strategy guiding the model to focus on easier negatives initially and progressively tackle more difficult ones. Secondly, it leverages the hierarchical and sequential information derived from the structural organization of statutes to evaluate the difficulty of samples. Lastly, it introduces a dynamic semantic difficulty assessment using the being-trained model itself, surpassing conventional static methods like BM25, adapting the negatives to the model's evolving competence. Experimental results on a real-world expert-annotated SAR dataset validate the effectiveness of CuSINeS across four different baselines, demonstrating its versatility.
翻译:本文提出了CuSINeS,一种用于提升法规条文检索(SAR)性能的负采样方法。CuSINeS包含三个关键贡献:首先,它采用基于课程的负采样策略,引导模型先关注简单负样本,再逐步处理困难负样本;其次,利用法规结构组织中的层次与序列信息评估样本难度;最后,引入基于训练模型自身的动态语义难度评估方法,超越BM25等传统静态方法,使负样本能适应模型不断进化的能力。在真实专家标注的SAR数据集上的实验结果验证了CuSINeS在四种不同基线模型上的有效性,展现了其广泛的适用性。