Multilingual proficiency presents a significant challenge for large language models (LLMs). English-centric models are usually suboptimal in other languages, particularly those that are linguistically distant from English. This performance discrepancy mainly stems from the imbalanced distribution of training data across languages during pre-training and instruction tuning stages. To address this problem, we propose a novel approach called CrossIn, which utilizes a mixed composition of cross-lingual instruction tuning data. Our method leverages the compressed representation shared by various languages to efficiently enhance the model's task-solving capabilities and multilingual proficiency within a single process. In addition, we introduce a multi-task and multi-faceted benchmark to evaluate the effectiveness of CrossIn. Experimental results demonstrate that our method substantially improves performance across tasks and languages, and we provide extensive insights into the impact of cross-lingual data volume and the integration of translation data on enhancing multilingual consistency and accuracy.
翻译:多语言能力对大语言模型(LLMs)构成了重大挑战。以英语为中心的模型在其他语言(尤其是与英语语言距离较远的语言)中通常表现欠佳。这种性能差异主要源于预训练和指令调优阶段训练数据在不同语言间分布不均衡。为解决此问题,我们提出了一种名为CrossIn的新方法,该方法利用跨语言指令调优数据的混合组合。我们的方法利用多种语言共享的压缩表示,在单一流程内高效提升模型的任务解决能力和多语言熟练度。此外,我们引入了一个多任务、多层面的基准来评估CrossIn的有效性。实验结果表明,我们的方法显著提升了跨任务和跨语言的性能,并深入分析了跨语言数据量以及翻译数据整合对提升多语言一致性和准确性的影响。