The adaption of multilingual pre-trained Large Language Models (LLMs) into eloquent and helpful assistants is essential to facilitate their use across different language regions. In that spirit, we are the first to conduct an extensive study of the performance of multilingual models on parallel, multi-turn instruction-tuning benchmarks across a selection of the most-spoken Indo-European languages. We systematically examine the effects of language and instruction dataset size on a mid-sized, multilingual LLM by instruction-tuning it on parallel instruction-tuning datasets. Our results demonstrate that instruction-tuning on parallel instead of monolingual corpora benefits cross-lingual instruction following capabilities by up to 4.6%. Furthermore, we show that the Superficial Alignment Hypothesis does not hold in general, as the investigated multilingual 7B parameter model presents a counter-example requiring large-scale instruction-tuning datasets. Finally, we conduct a human annotation study to understand the alignment between human-based and GPT-4-based evaluation within multilingual chat scenarios.
翻译:将多语言预训练大语言模型(LLMs)适配为健谈且有用的助手,对于促进其在各语言区域的应用至关重要。为此,我们首次开展了一项广泛研究,考察多语言模型在跨最具代表性的印欧语系语言的并行多轮指令微调基准上的表现。我们通过在一组并行指令微调数据集上对中型多语言LLM进行指令微调,系统性地研究了语言和指令数据集规模对其产生的影响。结果表明,相较于单语语料库,采用并行语料库进行指令微调可使跨语言指令遵循能力提升高达4.6%。此外,我们证明表面对齐假说(Superficial Alignment Hypothesis)在一般情况下并不成立——所研究的7B参数多语言模型呈现出一个反例,该模型需要大规模指令微调数据集。最后,我们通过人工标注研究,探讨了多语言对话场景中人类评估与基于GPT-4的评估之间的一致性。