The introduction of ChatGPT and the subsequent improvement of Large Language Models (LLMs) have prompted more and more individuals to turn to the use of ChatBots, both for information and assistance with decision-making. However, the information the user is after is often not formulated by these ChatBots objectively enough to be provided with a definite, globally accepted answer. Controversial topics, such as "religion", "gender identity", "freedom of speech", and "equality", among others, can be a source of conflict as partisan or biased answers can reinforce preconceived notions or promote disinformation. By exposing ChatGPT to such debatable questions, we aim to understand its level of awareness and if existing models are subject to socio-political and/or economic biases. We also aim to explore how AI-generated answers compare to human ones. For exploring this, we use a dataset of a social media platform created for the purpose of debating human-generated claims on polemic subjects among users, dubbed Kialo. Our results show that while previous versions of ChatGPT have had important issues with controversial topics, more recent versions of ChatGPT (gpt-3.5-turbo) are no longer manifesting significant explicit biases in several knowledge areas. In particular, it is well-moderated regarding economic aspects. However, it still maintains degrees of implicit libertarian leaning toward right-winged ideals which suggest the need for increased moderation from the socio-political point of view. In terms of domain knowledge on controversial topics, with the exception of the "Philosophical" category, ChatGPT is performing well in keeping up with the collective human level of knowledge. Finally, we see that sources of Bing AI have slightly more tendency to the center when compared to human answers. All the analyses we make are generalizable to other types of biases and domains.
翻译:ChatGPT的推出及随后大语言模型的改进,促使越来越多的人转向使用聊天机器人获取信息及辅助决策。然而,用户所需信息往往未被这些聊天机器人以足够客观的方式表述,从而无法提供明确且全球公认的答案。诸如“宗教”、“性别认同”、“言论自由”和“平等”等争议性话题可能成为冲突源头,因为带有党派立场或偏见的回答可能强化先入为主的观念或助长虚假信息。通过让ChatGPT面对此类争议性问题,我们旨在了解其认知水平,以及现有模型是否受到社会政治和/或经济偏见的影响。我们还旨在探索AI生成的回答与人类回答之间的差异。为此,我们使用了名为Kialo的社交媒体平台数据集,该平台旨在让用户就有争议性话题的人类论断进行辩论。结果显示,虽然早期版本的ChatGPT在争议性话题上存在重大问题,但较新版本(gpt-3.5-turbo)在多个知识领域已不再表现出显著的明确偏见。特别是在经济方面,其审核较为完善。然而,它仍隐含着一定程度的自由主义倾向,偏向右翼理念,这表明需要从社会政治角度加强审核。在争议性话题的领域知识方面,除“哲学”类别外,ChatGPT在保持与人类集体知识水平同步方面表现良好。最后,我们发现Bing AI的回答来源相较于人类回答,略微更倾向于中立。我们进行的分析均可推广至其他类型的偏见和领域。