Cross-lingual transfer in multilingual NLP has been widely explored in supervised fine-tuning contexts, where factors like data availability and linguistic similarity largely determine transfer quality. As the field shifts toward few-shot In-Context Learning (ICL), it is often presumed that insights from fine-tuning carry over unchanged. Yet this assumption has not been rigorously evaluated, leaving open the question of how to choose source languages for cross-lingual ICL. We conduct a broad empirical study of cross-lingual transfer in ICL spanning seven tasks, six models, and a typologically diverse set of languages. We further analyze language confusion, a key obstacle for generative tasks in cross-lingual ICL. Our results show that conventional fine-tuning-based expectations do not consistently apply in the ICL regime and point to alternative heuristics for selecting source languages effectively.
翻译:多语言自然语言处理中的跨语言迁移已在监督微调场景中得到广泛探索,其中数据可用性和语言相似性等因素在很大程度上决定了迁移质量。随着该领域转向少样本上下文学习(ICL),人们通常认为微调中的见解能够直接沿用。然而,这一假设尚未得到严格评估,导致跨语言ICL中如何选择源语言的问题悬而未决。我们针对跨语言ICL中的跨语言迁移进行了广泛的实证研究,涵盖七项任务、六个模型以及一组类型学多样的语言。我们进一步分析了语言混淆这一跨语言ICL生成任务中的关键障碍。研究结果表明,基于微调的传统预期在ICL场景中并非始终适用,并为有效选择源语言提供了替代性启发式方法。