Recent research has demonstrated that large pre-trained language models reflect societal biases expressed in natural language. The present paper introduces a simple method for probing language models to conduct a multilingual study of gender bias towards politicians. We quantify the usage of adjectives and verbs generated by language models surrounding the names of politicians as a function of their gender. To this end, we curate a dataset of 250k politicians worldwide, including their names and gender. Our study is conducted in seven languages across six different language modeling architectures. The results demonstrate that pre-trained language models' stance towards politicians varies strongly across analyzed languages. We find that while some words such as dead, and designated are associated with both male and female politicians, a few specific words such as beautiful and divorced are predominantly associated with female politicians. Finally, and contrary to previous findings, our study suggests that larger language models do not tend to be significantly more gender-biased than smaller ones.
翻译:近期研究已证明,大型预训练语言模型反映了自然语言中表达的社会偏见。本文提出一种简单方法,通过探针分析语言模型对政治人物进行多语言性别偏见研究。我们量化了语言模型在政治人物姓名周围产生的形容词和动词使用情况,并将其作为人物性别的函数进行分析。为此,我们整理了一个包含全球25万名政治人物姓名及性别的数据集。本研究涵盖六种不同语言建模架构的七种语言。结果表明,预训练语言模型对不同语言中政治人物的立场存在显著差异。我们发现,虽然"已故""指定"等词汇同时与男女政治人物关联,但"美丽""离异"等特定词汇主要与女性政治人物相关。最后,与先前研究结论相反,本研究表明较大规模的语言模型并不比小模型存在更显著的性别偏见。