In this paper, we present our solution to the Multilingual Information Retrieval Across a Continuum of Languages (MIRACL) challenge of WSDM CUP 2023\footnote{https://project-miracl.github.io/}. Our solution focuses on enhancing the ranking stage, where we fine-tune pre-trained multilingual transformer-based models with MIRACL dataset. Our model improvement is mainly achieved through diverse data engineering techniques, including the collection of additional relevant training data, data augmentation, and negative sampling. Our fine-tuned model effectively determines the semantic relevance between queries and documents, resulting in a significant improvement in the efficiency of the multilingual information retrieval process. Finally, Our team is pleased to achieve remarkable results in this challenging competition, securing 2nd place in the Surprise-Languages track with a score of 0.835 and 3rd place in the Known-Languages track with an average nDCG@10 score of 0.716 across the 16 known languages on the final leaderboard.
翻译:在本文中,我们介绍了针对WSDM CUP 2023多语言连续信息检索挑战赛(MIRACL)\footnote{https://project-miracl.github.io/}的解决方案。我们的方案聚焦于提升排序阶段,通过使用MIRACL数据集对预训练的多语言Transformer基础模型进行微调。模型改进主要得益于多样化的数据工程技术,包括收集额外相关训练数据、数据增强和负采样。经过微调的模型能够有效判定查询与文档之间的语义相关性,从而显著提升多语言信息检索过程的效率。最终,我们团队在该项充满挑战的竞赛中取得了令人瞩目的成绩:在惊喜语言赛道中以0.835的得分获得第二名,在已知语言赛道中以16种已知语言平均nDCG@10得分0.716的成绩位列第三。