As generative language models advance, users have started to utilize Large Language Models (LLMs) to assist in writing various types of content, including professional documents such as recommendation letters. Despite their convenience, these applications introduce unprecedented fairness concerns. As generated reference letters might be directly utilized by users in professional or academic scenarios, they have the potential to cause direct social harms, such as lowering success rates for female applicants. Therefore, it is imminent and necessary to comprehensively study fairness issues and associated harms in such real-world use cases for future mitigation and monitoring. In this paper, we critically examine gender bias in LLM-generated reference letters. Inspired by findings in social science, we design evaluation methods to manifest gender biases in LLM-generated letters through 2 dimensions: biases in language style and biases in lexical content. Furthermore, we investigate the extent of bias propagation by separately analyze bias amplification in model-hallucinated contents, which we define to be the hallucination bias of model-generated documents. Through benchmarking evaluation on 4 popular LLMs, including ChatGPT, Alpaca, Vicuna and StableLM, our study reveals significant gender biases in LLM-generated recommendation letters. Our findings further point towards the importance and imminence to recognize biases in LLM-generated professional documents.
翻译:随着生成式语言模型的发展,用户已开始利用大语言模型协助撰写各类内容,包括推荐信等专业文件。尽管此类应用带来便利,却也引发了前所未有的公平性问题。由于生成的推荐信可能被用户直接用于专业或学术场景,它们可能造成直接的社会危害,例如降低女性申请者的成功率。因此,亟需系统研究此类真实用例中的公平性问题及相关危害,以便未来制定缓解与监测措施。本文批判性地考察了LLM生成推荐信中的性别偏见。受社会科学研究成果启发,我们设计了评估方法,通过语言风格偏见和词汇内容偏见两个维度来揭示LLM生成信件中的性别偏见。进一步地,我们通过分别分析模型幻觉内容中的偏见放大现象(即模型生成文档的幻觉偏见)来研究偏见的传播程度。通过对包括ChatGPT、Alpaca、Vicuna和StableLM在内的四种主流LLM进行基准评估,本研究表明LLM生成的推荐信存在显著性别偏见。研究结果进一步揭示了识别LLM生成专业文档中偏见的紧迫性与重要性。