In this paper, we introduce the approach behind our submission for the MIRACL challenge, a WSDM 2023 Cup competition that centers on ad-hoc retrieval across 18 diverse languages. Our solution contains two neural-based models. The first model is a bi-encoder re-ranker, on which we apply a cross-lingual distillation technique to transfer ranking knowledge from English to the target language space. The second model is a cross-encoder re-ranker trained on multilingual retrieval data generated using neural machine translation. We further fine-tune both models using MIRACL training data and ensemble multiple rank lists to obtain the final result. According to the MIRACL leaderboard, our approach ranks 8th for the Test-A set and 2nd for the Test-B set among the 16 known languages.
翻译:本文介绍了我们针对MIRACL挑战赛的提交方案,该赛事是WSDM 2023杯赛,聚焦于18种不同语言的即席检索。我们的解决方案包含两个基于神经网络的模型。第一个模型是双编码器重排序器,我们对其应用跨语言蒸馏技术,将排序知识从英语迁移到目标语言空间。第二个模型是交叉编码器重排序器,使用基于神经机器翻译生成的多语检索数据进行训练。我们进一步利用MIRACL训练数据对两个模型进行微调,并集成多个排序列表以获得最终结果。根据MIRACL排行榜,在16个已知语言中,我们的方法在Test-A集中排名第8位,在Test-B集中排名第2位。