Large language models (LLMs) like ChatGPT have gained increasing prominence in artificial intelligence, making a profound impact on society and various industries like business and science. However, the presence of false information on the internet and in text corpus poses a significant risk to the reliability and safety of LLMs, underscoring the urgent need to understand the mechanisms of how false information impacts and spreads in LLMs. In this paper, we investigate how false information spreads in LLMs and affects related responses by conducting a series of experiments on the effects of source authority, injection paradigm, and information relevance. Specifically, we compare four authority levels of information sources (Twitter, web blogs, news reports, and research papers), two common knowledge injection paradigms (in-context injection and learning-based injection), and three degrees of information relevance (direct, indirect, and peripheral). The experimental results show that (1) False information will spread and contaminate related memories in LLMs via a semantic diffusion process, i.e., false information has global detrimental effects beyond its direct impact. (2) Current LLMs are susceptible to authority bias, i.e., LLMs are more likely to follow false information presented in a trustworthy style like news or research papers, which usually causes deeper and wider pollution of information. (3) Current LLMs are more sensitive to false information through in-context injection than through learning-based injection, which severely challenges the reliability and safety of LLMs even if all training data are trusty and correct. The above findings raise the need for new false information defense algorithms to address the global impact of false information, and new alignment algorithms to unbiasedly lead LLMs to follow internal human values rather than superficial patterns.
翻译:像ChatGPT这样的大型语言模型(LLMs)在人工智能领域日益突出,对商业、科学等各行各业及社会产生了深远影响。然而,互联网和文本语料库中虚假信息的存在,对LLMs的可靠性和安全性构成了重大风险,这凸显了理解虚假信息如何影响并在LLMs中传播机制的迫切性。本文通过一系列关于信息来源权威性、注入范式和信息相关性的实验,研究了虚假信息如何在LLMs中传播并影响相关回应。具体而言,我们比较了四种不同权威级别的信息来源(推特、网络博客、新闻报道和研究论文)、两种常见的知识注入范式(上下文内注入和基于学习的注入)以及三种不同相关程度的信息(直接相关、间接相关和边缘相关)。实验结果表明:(1)虚假信息会通过语义扩散过程在LLMs中传播并污染相关记忆,即虚假信息会产生超出其直接影响范围的全局性有害效应。(2)当前LLMs易受权威偏见影响,即它们更倾向于遵循以新闻报道或研究论文等可信风格呈现的虚假信息,这通常会导致信息污染得更深更广。(3)与基于学习的注入相比,当前LLMs对通过上下文内注入的虚假信息更为敏感,这严重挑战了LLMs的可靠性和安全性——即便所有训练数据都真实正确。上述发现表明,需要开发新型虚假信息防御算法来应对其全局影响,以及新的对齐算法来无偏见地引导LLMs遵循人类内在价值观而非表面模式。