This work is intended as a voice in the discussion over previous claims that a pretrained large language model (LLM) based on the Transformer model architecture can be sentient. Such claims have been made concerning the LaMDA model and also concerning the current wave of LLM-powered chatbots, such as ChatGPT. This claim, if confirmed, would have serious ramifications in the Natural Language Processing (NLP) community due to wide-spread use of similar models. However, here we take the position that such a large language model cannot be sentient, or conscious, and that LaMDA in particular exhibits no advances over other similar models that would qualify it. We justify this by analysing the Transformer architecture through Integrated Information Theory of consciousness. We see the claims of sentience as part of a wider tendency to use anthropomorphic language in NLP reporting. Regardless of the veracity of the claims, we consider this an opportune moment to take stock of progress in language modelling and consider the ethical implications of the task. In order to make this work helpful for readers outside the NLP community, we also present the necessary background in language modelling.
翻译:本文旨在回应先前关于基于Transformer模型架构的预训练大型语言模型(LLM)具有感知能力的主张。这些主张涉及LaMDA模型以及当前基于LLM的聊天机器人(如ChatGPT)浪潮。若该主张得到证实,将因类似模型的广泛应用而对自然语言处理(NLP)领域产生严重影响。然而,本文认为此类大型语言模型无法具备感知能力或意识,特别是LaMDA模型相较于其他类似模型并未展现出使其具备感知资格的任何进步。我们通过意识整合信息理论分析Transformer架构来论证这一观点。我们将这些感知主张视为NLP报道中更广泛使用拟人化语言倾向的一部分。无论这些主张真实性如何,我们认为此刻正是审视语言模型进展、并思考该任务伦理影响的良机。为便于NLP领域外读者理解,文中还介绍了语言建模的必要背景知识。