Little attention is placed on analyzing nationality bias in language models, especially when nationality is highly used as a factor in increasing the performance of social NLP models. This paper examines how a text generation model, GPT-2, accentuates pre-existing societal biases about country-based demonyms. We generate stories using GPT-2 for various nationalities and use sensitivity analysis to explore how the number of internet users and the country's economic status impacts the sentiment of the stories. To reduce the propagation of biases through large language models (LLM), we explore the debiasing method of adversarial triggering. Our results show that GPT-2 demonstrates significant bias against countries with lower internet users, and adversarial triggering effectively reduces the same.
翻译:在语言模型中,国籍偏见分析鲜受关注,尤其当国籍作为提升社会性自然语言处理(NLP)模型性能的重要因素时。本文考察了文本生成模型GPT-2如何强化关于国家籍民词(demonyns)的既有社会偏见。我们利用GPT-2针对多种国籍生成故事,并通过敏感性分析探究互联网用户数量及国家经济状况对故事情感倾向的影响。为减少大型语言模型(LLM)中偏见的传播,我们探索了基于对抗触发(adversarial triggering)的去偏方法。结果表明,GPT-2对互联网用户较少的国家表现出显著偏见,而对抗触发能有效降低此类偏见。