We consider the questions of whether or not large language models (LLMs) have beliefs, and, if they do, how we might measure them. First, we evaluate two existing approaches, one due to Azaria and Mitchell (2023) and the other to Burns et al. (2022). We provide empirical results that show that these methods fail to generalize in very basic ways. We then argue that, even if LLMs have beliefs, these methods are unlikely to be successful for conceptual reasons. Thus, there is still no lie-detector for LLMs. After describing our empirical results we take a step back and consider whether or not we should expect LLMs to have something like beliefs in the first place. We consider some recent arguments aiming to show that LLMs cannot have beliefs. We show that these arguments are misguided. We provide a more productive framing of questions surrounding the status of beliefs in LLMs, and highlight the empirical nature of the problem. We conclude by suggesting some concrete paths for future work.
翻译:我们探讨了大语言模型(LLMs)是否具有信念,以及如果具有信念,应如何度量这两个问题。首先,我们评估了Azaria与Mitchell(2023)及Burns等人(2022)提出的两种现有方法,并通过实证结果表明这些方法在最基本层面无法泛化。接着我们论证,即便LLMs具备信念,由于概念层面的原因,这些方法也难以奏效——因此,当前仍不存在针对LLMs的谎言检测器。在呈现实证结果后,我们退而思考:是否应首先期望LLMs具备类似信念的特性?我们考察了近期主张LLMs不可能具有信念的若干论证,指出这些论证存在偏差。通过构建关于LLMs信念状态问题的更具建设性框架,我们凸显了该问题的实证本质。最后,我们为未来研究提出了若干具体方向。