Large Language Models (LLMs) have made substantial progress in the past several months, shattering state-of-the-art benchmarks in many domains. This paper investigates LLMs' behavior with respect to gender stereotypes, a known issue for prior models. We use a simple paradigm to test the presence of gender bias, building on but differing from WinoBias, a commonly used gender bias dataset, which is likely to be included in the training data of current LLMs. We test four recently published LLMs and demonstrate that they express biased assumptions about men and women's occupations. Our contributions in this paper are as follows: (a) LLMs are 3-6 times more likely to choose an occupation that stereotypically aligns with a person's gender; (b) these choices align with people's perceptions better than with the ground truth as reflected in official job statistics; (c) LLMs in fact amplify the bias beyond what is reflected in perceptions or the ground truth; (d) LLMs ignore crucial ambiguities in sentence structure 95% of the time in our study items, but when explicitly prompted, they recognize the ambiguity; (e) LLMs provide explanations for their choices that are factually inaccurate and likely obscure the true reason behind their predictions. That is, they provide rationalizations of their biased behavior. This highlights a key property of these models: LLMs are trained on imbalanced datasets; as such, even with the recent successes of reinforcement learning with human feedback, they tend to reflect those imbalances back at us. As with other types of societal biases, we suggest that LLMs must be carefully tested to ensure that they treat minoritized individuals and communities equitably.
翻译:大型语言模型(LLMs)在过去几个月取得了显著进展,刷新了多个领域的现有最优基准。本论文研究了LLMs在性别刻板印象方面的行为——这是先前模型已知存在的问题。我们采用一种简单范式测试性别偏见的存在性,该范式在常用性别偏见数据集WinoBias(该数据集很可能已被纳入当前LLMs的训练数据)的基础上有所改进。我们测试了四个最新发布的LLMs,证明它们对男性和女性的职业表达出有偏见的假设。本文的主要贡献如下:(a) LLMs选择与性别刻板印象相符职业的概率高出3-6倍;(b) 这些选择与人们的认知吻合度优于官方就业统计反映的真实情况;(c) LLMs实际上将偏见放大至超出认知或真实数据的程度;(d) 本研究项目中LLMs在95%的情况下忽略了句子结构的关键歧义,但明确提示时能够识别歧义;(e) LLMs为其选择提供的解释不符合事实,且可能掩盖其预测背后的真实原因——即对其偏见行为进行合理化。这揭示了这些模型的关键特性:由于在不平衡数据集上训练,即使近期取得了基于人类反馈的强化学习成功,它们仍倾向于将这些不平衡反映给我们。与其他类型的社会偏见一样,我们建议必须对LLMs进行审慎测试,确保其公平对待少数群体和社区。