As large language models (LLMs) advance to produce human-like arguments in some contexts, the number of settings applicable for human-AI collaboration broadens. Specifically, we focus on subjective decision-making, where a decision is contextual, open to interpretation, and based on one's beliefs and values. In such cases, having multiple arguments and perspectives might be particularly useful for the decision-maker. Using subtle sexism online as an understudied application of subjective decision-making, we suggest that LLM output could effectively provide diverse argumentation to enrich subjective human decision-making. To evaluate the applicability of this case, we conducted an interview study (N=20) where participants evaluated the perceived authorship, relevance, convincingness, and trustworthiness of human and AI-generated explanation-text, generated in response to instances of subtle sexism from the internet. In this workshop paper, we focus on one troubling trend in our results related to opinions and experiences displayed in LLM argumentation. We found that participants rated explanations that contained these characteristics as more convincing and trustworthy, particularly so when those opinions and experiences aligned with their own opinions and experiences. We describe our findings, discuss the troubling role that confirmation bias plays, and bring attention to the ethical challenges surrounding the AI generation of human-like experiences.
翻译:随着大型语言模型(LLM)在某些情境下生成类人论证的能力不断进步,人机协作的适用场景范围也随之拓展。我们重点关注主观决策——这类决策具有情境依赖性、可解读空间大,且基于个人的信念与价值观。在此类情境中,多元的论证与视角对决策者尤为有益。以网络隐性性别歧视为例(这一主观决策应用领域此前研究不足),我们认为LLM输出能够有效提供多样化论证,从而丰富人类的主观决策过程。为评估该场景的适用性,我们开展了访谈研究(N=20),让参与者评估由人类与AI生成的对网络隐性性别歧视案例的回应文本,并从作者感知、相关性、说服力及可信度四个维度进行评价。在本研讨会论文中,我们聚焦于LLM论证中涉及意见与个人经历的一项令人担忧的趋势。研究发现,当解释文本包含此类特征时,参与者对其评价更高(更具说服力与可信度),尤其当这些意见与经历与参与者自身的意见和经历相契合时更为显著。我们描述了研究发现,讨论了确认偏误在此过程中扮演的令人担忧的角色,并揭示了AI生成类人经历所引发的伦理挑战。