Language serves as a powerful tool for the manifestation of societal belief systems. In doing so, it also perpetuates the prevalent biases in our society. Gender bias is one of the most pervasive biases in our society and is seen in online and offline discourses. With LLMs increasingly gaining human-like fluency in text generation, gaining a nuanced understanding of the biases these systems can generate is imperative. Prior work often treats gender bias as a binary classification task. However, acknowledging that bias must be perceived at a relative scale; we investigate the generation and consequent receptivity of manual annotators to bias of varying degrees. Specifically, we create the first dataset of GPT-generated English text with normative ratings of gender bias. Ratings were obtained using Best--Worst Scaling -- an efficient comparative annotation framework. Next, we systematically analyze the variation of themes of gender biases in the observed ranking and show that identity-attack is most closely related to gender bias. Finally, we show the performance of existing automated models trained on related concepts on our dataset.
翻译:语言作为社会信念体系展现的有力工具,同时也固化了社会中普遍存在的偏见。性别偏见是当今社会中最普遍的偏见形式之一,在线上线下话语中均有体现。随着大型语言模型在文本生成方面日益接近人类般的流畅度,深入理解这些系统可能生成的偏见变得至关重要。以往研究常将性别偏见视为二元分类任务,但认识到偏见需在相对尺度上感知后,我们探究了不同强度的偏见生成及其对人工标注者的接受度。具体而言,我们创建了首个附带性别偏见规范性评级的GPT生成英文文本数据集。采用最佳-最差缩放法(一种高效的比较性标注框架)获取评级数据。继而,我们系统分析了观察到的排序中性别偏见主题的变化,发现身份攻击与性别偏见的关联最为密切。最终,我们展示了基于相关概念训练的现有自动化模型在该数据集上的表现。