Little attention is placed on analyzing nationality bias in language models, especially when nationality is highly used as a factor in increasing the performance of social NLP models. This paper examines how a text generation model, GPT-2, accentuates pre-existing societal biases about country-based demonyms. We generate stories using GPT-2 for various nationalities and use sensitivity analysis to explore how the number of internet users and the country's economic status impacts the sentiment of the stories. To reduce the propagation of biases through large language models (LLM), we explore the debiasing method of adversarial triggering. Our results show that GPT-2 demonstrates significant bias against countries with lower internet users, and adversarial triggering effectively reduces the same.
翻译:在语言模型中,关于国籍偏见的分析很少受到关注,尤其是当国籍被广泛用作提升社会性NLP模型性能的一个因素时。本文研究了文本生成模型GPT-2如何加剧关于国家居民名称的现有社会偏见。我们使用GPT-2为多种国籍生成故事,并通过敏感性分析探讨互联网用户数量及国家经济状况对故事情感倾向的影响。为减少通过大型语言模型(LLM)传播的偏见,我们探索了使用对抗触发(adversarial triggering)的去偏方法。结果表明,GPT-2对互联网用户数量较少的国家表现出显著偏见,而对抗触发有效减少了这种偏见。