Large language models are quickly gaining momentum, yet are found to demonstrate gender bias in their responses. In this paper, we conducted a content analysis of social media discussions to gauge public perceptions of gender bias in LLMs which are trained in different cultural contexts, i.e., ChatGPT, a US-based LLM, or Ernie, a China-based LLM. People shared both observations of gender bias in their personal use and scientific findings about gender bias in LLMs. A difference between the two LLMs was seen -- ChatGPT was more often found to carry implicit gender bias, e.g., associating men and women with different profession titles, while explicit gender bias was found in Ernie's responses, e.g., overly promoting women's pursuit of marriage over career. Based on the findings, we reflect on the impact of culture on gender bias and propose governance recommendations to regulate gender bias in LLMs.
翻译:大型语言模型正迅速获得发展势头,但在其回复中被发现存在性别偏见。本文通过对社交媒体讨论的内容分析,评估公众对在不同文化背景下训练的大型语言模型(即基于美国的ChatGPT与基于中国的文心一言)中性别偏见的认知。人们既分享了个人使用中对性别偏见的观察,也分享了关于大型语言模型中性别偏见的科学发现。两种大型语言模型之间存在差异——ChatGPT更常被发现携带隐性性别偏见,例如将男性和女性与不同职业头衔相关联;而文心一言的回复中则发现了显性性别偏见,例如过度提倡女性追求婚姻而非事业。基于这些发现,我们反思文化对性别偏见的影响,并提出治理建议以规范大型语言模型中的性别偏见。