As 3rd-person pronoun usage shifts to include novel forms, e.g., neopronouns, we need more research on identity-inclusive NLP. Exclusion is particularly harmful in one of the most popular NLP applications, machine translation (MT). Wrong pronoun translations can discriminate against marginalized groups, e.g., non-binary individuals (Dev et al., 2021). In this ``reality check'', we study how three commercial MT systems translate 3rd-person pronouns. Concretely, we compare the translations of gendered vs. gender-neutral pronouns from English to five other languages (Danish, Farsi, French, German, Italian), and vice versa, from Danish to English. Our error analysis shows that the presence of a gender-neutral pronoun often leads to grammatical and semantic translation errors. Similarly, gender neutrality is often not preserved. By surveying the opinions of affected native speakers from diverse languages, we provide recommendations to address the issue in future MT research.
翻译:随着第三人称代词的使用扩展到包含新形式(如新代词),我们需要更多关于包容性自然语言处理的研究。在机器翻译这一最流行的自然语言处理应用中,排斥问题尤为严重。错误的代词翻译可能歧视边缘化群体,例如非二元性别者(Dev et al., 2021)。在这项"现实检验"中,我们研究了三种商业机器翻译系统如何处理第三人称代词。具体而言,我们比较了英语到五种其他语言(丹麦语、波斯语、法语、德语、意大利语)以及丹麦语到英语的翻译中,有性别代词与性别中立代词的差异。我们的错误分析表明,性别中立代词的出现常导致语法和语义翻译错误。同样,性别中立性也常常无法保留。通过调查来自不同语言的受影响母语者的观点,我们为未来机器翻译研究如何解决这一问题提供了建议。