This paper reports on an audit study of generative AI systems (ChatGPT, Bing Chat, and Perplexity) which investigates how these new search engines construct responses and establish authority for topics of public importance. We collected system responses using a set of 48 authentic queries for 4 topics over a 7-day period and analyzed the data using sentiment analysis, inductive coding and source classification. Results provide an overview of the nature of system responses across these systems and provide evidence of sentiment bias based on the queries and topics, and commercial and geographic bias in sources. The quality of sources used to support claims is uneven, relying heavily on News and Media, Business and Digital Media websites. Implications for system users emphasize the need to critically examine Generative AI system outputs when making decisions related to public interest and personal well-being.
翻译:本文报告了一项关于生成式AI系统(ChatGPT、Bing Chat和Perplexity)的审计研究,旨在探究这些新型搜索引擎如何构建回答并为重要公共议题建立权威性。我们在7天内针对4个主题收集了48个真实查询的系统响应,并采用情感分析、归纳编码和来源分类方法对数据进行分析。结果呈现了各系统响应的整体特征,证明了基于查询和主题的情感偏见,以及来源中存在的商业与地域偏见。用于支撑论断的资料来源质量参差不齐,高度依赖新闻媒体、商业及数字媒体网站。研究对系统用户的启示强调,在做出涉及公共利益与个人福祉的决策时,必须批判性地审视生成式AI系统的输出内容。