Multilingual generative language models (LMs) are increasingly fluent in a large variety of languages. Trained on the concatenation of corpora in multiple languages, they enable powerful transfer from high-resource languages to low-resource ones. However, it is still unknown what cultural biases are induced in the predictions of these models. In this work, we focus on one language property highly influenced by culture: formality. We analyze the formality distributions of XGLM and BLOOM's predictions, two popular generative multilingual language models, in 5 languages. We classify 1,200 generations per language as formal, informal, or incohesive and measure the impact of the prompt formality on the predictions. Overall, we observe a diversity of behaviors across the models and languages. For instance, XGLM generates informal text in Arabic and Bengali when conditioned with informal prompts, much more than BLOOM. In addition, even though both models are highly biased toward the formal style when prompted neutrally, we find that the models generate a significant amount of informal predictions even when prompted with formal text. We release with this work 6,000 annotated samples, paving the way for future work on the formality of generative multilingual LMs.
翻译:多语言生成式语言模型(LMs)在大范围语言中的流畅度日益提升。这些模型通过多语言语料库的拼接训练,能够实现从高资源语言到低资源语言的强大迁移能力。然而,这些模型预测中蕴含的文化偏见仍属未知。本研究聚焦于一种深受文化影响的语言属性:形式性。我们分析了两种主流多语言生成式模型——XGLM和BLOOM——在5种语言中的形式性分布。针对每种语言,我们对1200个生成文本进行分类(正式、非正式或缺乏连贯性),并测量提示形式性对预测结果的影响。总体而言,我们观察到模型与语言之间的行为多样性。例如,在非正式提示条件下,XGLM生成的阿拉伯语和孟加拉语非正式文本显著多于BLOOM。此外,尽管在中性提示下两种模型均高度偏向正式风格,但即使使用正式提示,模型仍会产生大量非正式预测。本研究公开了6000个标注样本,为后续多语言生成式LM形式性研究奠定基础。