Large, multilingual language models exhibit surprisingly good zero- or few-shot machine translation capabilities, despite having never seen the intentionally-included translation examples provided to typical neural translation systems. We investigate the role of incidental bilingualism -- the unintentional consumption of bilingual signals, including translation examples -- in explaining the translation capabilities of large language models, taking the Pathways Language Model (PaLM) as a case study. We introduce a mixed-method approach to measure and understand incidental bilingualism at scale. We show that PaLM is exposed to over 30 million translation pairs across at least 44 languages. Furthermore, the amount of incidental bilingual content is highly correlated with the amount of monolingual in-language content for non-English languages. We relate incidental bilingual content to zero-shot prompts and show that it can be used to mine new prompts to improve PaLM's out-of-English zero-shot translation quality. Finally, in a series of small-scale ablations, we show that its presence has a substantial impact on translation capabilities, although this impact diminishes with model scale.
翻译:大型多语言语言模型展现出令人惊讶的零样本或少样本机器翻译能力,尽管它们从未接触过传统神经翻译系统中刻意包含的翻译示例。本研究以Pathways语言模型(PaLM)为案例,探究偶然双语现象——即无意中摄入包含翻译示例在内的双语信号——在解释大型语言模型翻译能力中的作用。我们引入了一种混合方法,用于大规模度量并理解偶然双语现象。研究表明,PaLM在至少44种语言中接触了超过3000万个翻译对。此外,非英语语言的偶然双语内容量与其单语内容量高度相关。我们将偶然双语内容与零样本提示相关联,并证明可利用其挖掘新提示来提升PaLM从英语向外的零样本翻译质量。最后,通过一系列小规模消融实验,我们证实偶然双语现象的存在对翻译能力有显著影响,但该影响随模型规模扩大而减弱。