This work is intended as a voice in the discussion over previous claims that a pretrained large language model (LLM) based on the Transformer model architecture can be sentient. Such claims have been made concerning the LaMDA model and also concerning the current wave of LLM-powered chatbots, such as ChatGPT. This claim, if confirmed, would have serious ramifications in the Natural Language Processing (NLP) community due to wide-spread use of similar models. However, here we take the position that such a large language model cannot be sentient, or conscious, and that LaMDA in particular exhibits no advances over other similar models that would qualify it. We justify this by analysing the Transformer architecture through Integrated Information Theory of consciousness. We see the claims of sentience as part of a wider tendency to use anthropomorphic language in NLP reporting. Regardless of the veracity of the claims, we consider this an opportune moment to take stock of progress in language modelling and consider the ethical implications of the task. In order to make this work helpful for readers outside the NLP community, we also present the necessary background in language modelling.
翻译:本文旨在回应此前关于基于Transformer模型架构的预训练大型语言模型具有感知能力的主张。这些主张涉及LaMDA模型以及当前由大型语言模型驱动的聊天机器人(如ChatGPT)浪潮。若该主张被证实,由于类似模型的广泛应用,将对自然语言处理学界产生深远影响。然而,本文持相反立场:此类大型语言模型不具备感知能力或意识,尤其LaMDA模型相较于其他类似模型并未展现任何足以证明其感知能力的突破性进展。我们通过整合信息理论对Transformer架构进行分析来论证这一观点。我们将感知能力主张视为自然语言处理领域报告中拟人化语言泛化趋势的体现。无论该主张是否属实,我们认为此时正是审视语言建模进展、思考该任务伦理影响的恰当时机。为使研究成果惠及自然语言处理领域外的读者,本文亦提供了语言建模的必要背景知识。