Scaling up neural models has yielded significant advancements in a wide array of tasks, particularly in language generation. Previous studies have found that the performance of neural models frequently adheres to predictable scaling laws, correlated with factors such as training set size and model size. This insight is invaluable, especially as large-scale experiments grow increasingly resource-intensive. Yet, such scaling law has not been fully explored in dense retrieval due to the discrete nature of retrieval metrics and complex relationships between training data and model sizes in retrieval tasks. In this study, we investigate whether the performance of dense retrieval models follows the scaling law as other neural models. We propose to use contrastive log-likelihood as the evaluation metric and conduct extensive experiments with dense retrieval models implemented with different numbers of parameters and trained with different amounts of annotated data. Results indicate that, under our settings, the performance of dense retrieval models follows a precise power-law scaling related to the model size and the number of annotations. Additionally, we examine scaling with prevalent data augmentation methods to assess the impact of annotation quality, and apply the scaling law to find the best resource allocation strategy under a budget constraint. We believe that these insights will significantly contribute to understanding the scaling effect of dense retrieval models and offer meaningful guidance for future research endeavors.
翻译:扩展神经模型已经在众多任务中取得了显著进展,尤其是在语言生成领域。先前研究发现,神经模型的性能通常遵循可预测的规模定律,这与训练集大小和模型大小等因素相关。这一洞察极具价值,尤其是在大规模实验资源消耗日益增长的背景下。然而,由于检索指标的离散特性以及检索任务中训练数据与模型大小之间的复杂关系,密集检索领域的规模定律尚未得到充分探索。本研究旨在探究密集检索模型的性能是否像其他神经模型一样遵循规模定律。我们提出使用对比对数似然作为评估指标,并基于不同参数数量和不同规模标注数据训练的密集检索模型开展了大量实验。结果表明,在我们的设定下,密集检索模型性能严格遵循与模型大小和标注数量相关的幂律规模定律。此外,我们通过主流数据增强方法考察了规模效应以评估标注质量的影响,并应用该规模定律在预算约束下寻找最优资源分配策略。我们相信,这些发现将有助于深入理解密集检索模型的规模效应,并为未来研究提供有价值的指导。