Pre-trained multilingual language models (PMLMs) are commonly used when dealing with data from multiple languages and cross-lingual transfer. However, PMLMs are trained on varying amounts of data for each language. In practice this means their performance is often much better on English than many other languages. We explore to what extent this also applies to moral norms. Do the models capture moral norms from English and impose them on other languages? Do the models exhibit random and thus potentially harmful beliefs in certain languages? Both these issues could negatively impact cross-lingual transfer and potentially lead to harmful outcomes. In this paper, we (1) apply the MoralDirection framework to multilingual models, comparing results in German, Czech, Arabic, Chinese, and English, (2) analyse model behaviour on filtered parallel subtitles corpora, and (3) apply the models to a Moral Foundations Questionnaire, comparing with human responses from different countries. Our experiments demonstrate that, indeed, PMLMs encode differing moral biases, but these do not necessarily correspond to cultural differences or commonalities in human opinions. We release our code and models.
翻译:预训练多语言语言模型(PMLMs)在处理多语言数据和跨语言迁移时被广泛使用。然而,PMLMs针对每种语言训练的数据量不同,实际应用中它们的性能通常在英语上远优于其他语言。我们探究了这种情况在道德规范层面的适用程度:模型是否会从英语中捕获道德规范并将其强加于其他语言?模型是否在某些语言中表现出随机的、可能有害的信念?这两个问题都可能对跨语言迁移产生负面影响,并可能导致有害后果。本文中,我们(1)将MoralDirection框架应用于多语言模型,对比德语、捷克语、阿拉伯语、中文和英语的结果;(2)分析模型在过滤后的平行字幕语料库上的行为;以及(3)将模型应用于道德基础问卷,并与来自不同国家的人类回答进行对比。我们的实验表明,确实,PMLMs编码了不同的道德偏见,但这些偏见并不一定对应人类观点中的文化差异或共性。我们已公开发布代码和模型。