We present GEST -- a new dataset for measuring gender-stereotypical reasoning in masked LMs and English-to-X machine translation systems. GEST contains samples that are compatible with 9 Slavic languages and English for 16 gender stereotypes about men and women (e.g., Women are beautiful, Men are leaders). The definition of said stereotypes was informed by gender experts. We used GEST to evaluate 11 masked LMs and 4 machine translation systems. We discovered significant and consistent amounts of stereotypical reasoning in almost all the evaluated models and languages.
翻译:我们提出GEST——一个用于评估遮蔽语言模型和英语至X语言机器翻译系统中性别刻板推理的新数据集。GEST包含与9种斯拉夫语言和英语兼容的样本,涵盖16个关于男性和女性的性别刻板印象(例如,女性是美丽的,男性是领导者)。这些刻板印象的定义由性别专家提供。我们使用GEST评估了11个遮蔽语言模型和4个机器翻译系统。研究发现在几乎所有被评估的模型和语言中,均存在显著且一致的刻板推理现象。