Pre-trained language models have been known to perpetuate biases from the underlying datasets to downstream tasks. However, these findings are predominantly based on monolingual language models for English, whereas there are few investigative studies of biases encoded in language models for languages beyond English. In this paper, we fill this gap by analysing gender bias in West Slavic language models. We introduce the first template-based dataset in Czech, Polish, and Slovak for measuring gender bias towards male, female and non-binary subjects. We complete the sentences using both mono- and multilingual language models and assess their suitability for the masked language modelling objective. Next, we measure gender bias encoded in West Slavic language models by quantifying the toxicity and genderness of the generated words. We find that these language models produce hurtful completions that depend on the subject's gender. Perhaps surprisingly, Czech, Slovak, and Polish language models produce more hurtful completions with men as subjects, which, upon inspection, we find is due to completions being related to violence, death, and sickness.
翻译:预训练语言模型已知会将底层数据集中的偏见传播至下游任务。然而,这些发现主要基于英语的单语言模型,对于英语以外语言编码的偏见研究仍较为匮乏。本文通过分析西斯拉夫语言模型中的性别偏见来填补这一空白。我们首次构建了基于模板的捷克语、波兰语和斯洛伐克语数据集,用于测量针对男性、女性和非二元主体的性别偏见。我们使用单语言和多语言语言模型完成句子填空,并评估它们对掩码语言建模目标的适用性。接着,通过量化生成词汇的毒性程度和性别属性,我们测量了西斯拉夫语言模型编码的性别偏见。研究发现,这些语言模型会产生有害的补全结果,且这些结果依赖于主体的性别。出人意料的是,捷克语、斯洛伐克语和波兰语模型在以男性为主体时会产生更多有害补全,进一步分析表明,这源于补全结果与暴力、死亡和疾病相关的内容。