Large Language Models (LLMs) have recently emerged as an effective tool to assist individuals in writing various types of content, including professional documents such as recommendation letters. Though bringing convenience, this application also introduces unprecedented fairness concerns. Model-generated reference letters might be directly used by users in professional scenarios. If underlying biases exist in these model-constructed letters, using them without scrutinization could lead to direct societal harms, such as sabotaging application success rates for female applicants. In light of this pressing issue, it is imminent and necessary to comprehensively study fairness issues and associated harms in this real-world use case. In this paper, we critically examine gender biases in LLM-generated reference letters. Drawing inspiration from social science findings, we design evaluation methods to manifest biases through 2 dimensions: (1) biases in language style and (2) biases in lexical content. We further investigate the extent of bias propagation by analyzing the hallucination bias of models, a term that we define to be bias exacerbation in model-hallucinated contents. Through benchmarking evaluation on 2 popular LLMs- ChatGPT and Alpaca, we reveal significant gender biases in LLM-generated recommendation letters. Our findings not only warn against using LLMs for this application without scrutinization, but also illuminate the importance of thoroughly studying hidden biases and harms in LLM-generated professional documents.
翻译:大型语言模型(LLMs)近期已成为协助个人撰写各类内容(包括推荐信等专业文档)的有效工具。尽管带来了便利,但这一应用也引发了前所未有的公平性问题。模型生成的推荐信可能会被用户直接用于专业场景。如果这些模型构建的信件中存在潜在偏见,未经审视地使用它们可能导致直接的社会危害,例如降低女性申请者的申请成功率。鉴于这一紧迫问题,亟需且有必要全面研究这一真实用例中的公平性问题和相关危害。本文批判性地审视了LLM生成推荐信中的性别偏见。受社会科学研究发现的启发,我们设计了通过两个维度体现偏见的评估方法:(1)语言风格中的偏见,(2)词汇内容中的偏见。我们进一步通过分析模型的幻觉偏见——即我们定义为模型幻觉内容中偏见加剧的现象——来探究偏见传播的程度。通过对两个流行LLM(ChatGPT和Alpaca)的基准评估,我们揭示了LLM生成推荐信中显著的性别偏见。我们的发现不仅警示人们在使用LLM进行此类应用时需加以审视,也阐明了深入研究LLM生成专业文档中隐藏偏见与危害的重要性。