With the increasing utilization of multilingual text information, Cross-Lingual Information Retrieval (CLIR) has become a crucial research area. However, the impact of training data composition on both CLIR and Mono-Lingual Information Retrieval (IR) performance remains under-explored. To systematically investigate this data-centric aspect, we construct linguistically parallel Korean-English datasets and train retrieval models with various language combinations. Our experiments reveal that the language composition of training data significantly influences IR performance, exhibiting important inter-lingual correlations: CLIR performance improves with specific language pairs, while Mono-Lingual IR performance declines. Our work demonstrates that Model Merging can effectively mitigate this trade-off, achieving strong CLIR results while preserving Mono-Lingual IR capabilities. Our findings underscore the effects of linguistic configuration of training data on both CLIR and Mono-Lingual IR, and present Model Merging as a viable strategy to optimize performance across these tasks.
翻译:随着多语言文本信息的广泛应用,跨语言信息检索(CLIR)已成为重要的研究领域。然而,训练数据的构成对CLIR和单语言信息检索(IR)性能的影响仍未得到充分探索。为系统研究这一数据驱动因素,我们构建了语言平行的韩英数据集,并采用多种语言组合训练检索模型。实验结果表明,训练数据的语言构成显著影响IR性能,并呈现出重要的跨语言关联:特定语言对能提升CLIR性能,而单语言IR性能则有所下降。我们的研究证明,模型融合可以有效缓解这一权衡,在保持单语言IR能力的同时,实现优异的CLIR效果。研究结果强调了训练数据语言配置对CLIR和单语言IR的共同影响,并提出了模型融合作为优化跨任务性能的可行策略。