With the introduction of ChatGPT, OpenAI made large language models (LLM) accessible to users with limited IT expertise. However, users with no background in natural language processing (NLP) might lack a proper understanding of LLMs. Thus the awareness of their inherent limitations, and therefore will take the systems' output at face value. In this paper, we systematically analyse prompts and the generated responses to identify possible problematic issues with a special focus on gender biases, which users need to be aware of when processing the system's output. We explore how ChatGPT reacts in English and German if prompted to answer from a female, male, or neutral perspective. In an in-depth investigation, we examine selected prompts and analyse to what extent responses differ if the system is prompted several times in an identical way. On this basis, we show that ChatGPT is indeed useful for helping non-IT users draft texts for their daily work. However, it is absolutely crucial to thoroughly check the system's responses for biases as well as for syntactic and grammatical mistakes.
翻译:随着ChatGPT的推出,OpenAI使缺乏IT专业知识的用户也能够使用大型语言模型(LLM)。然而,没有自然语言处理(NLP)背景的用户可能缺乏对LLM的充分理解,因而意识不到其固有限制,从而会轻信系统输出的表面内容。在本文中,我们系统地分析了提示词及其生成的回复,以识别可能存在的问题,尤其关注性别偏见——用户在处理系统输出时需要警惕的问题。我们探讨了当以女性、男性或中性视角进行提示时,ChatGPT在英语和德语中的反应。通过一项深入调查,我们检查了选定的提示词,并分析当系统以相同方式被多次提示时,回复在多大程度上存在差异。在此基础上,我们表明ChatGPT确实有助于帮助非IT用户起草日常工作中的文本。然而,彻底检查系统回复中是否存在偏见以及句法和语法错误至关重要。