Artificial Intelligence (AI) has been rapidly advancing and has demonstrated its ability to perform a wide range of cognitive tasks, including language processing, visual recognition, and decision-making. Part of this progress is due to LLMs (Large Language Models) like those of the GPT (Generative Pre-Trained Transformers) family. These models are capable of exhibiting behavior that can be perceived as intelligent. Most authors in Neuropsychology consider intelligent behavior to depend on a number of overarching skills, or Executive Functions (EFs), which rely on the correct functioning of neural networks in the frontal lobes, and have developed a series of tests to evaluate them. In this work, we raise the question of whether LLMs are developing executive functions similar to those of humans as part of their learning, and we evaluate the planning function and working memory of GPT using the popular Towers of Hanoi method. Additionally, we introduce a new variant of the classical method in order to avoid that the solutions are found in the LLM training data (dataleakeage). Preliminary results show that LLMs generates near-optimal solutions in Towers of Hanoi related tasks, adheres to task constraints, and exhibits rapid planning capabilities and efficient working memory usage, indicating a potential development of executive functions. However, these abilities are quite limited and worse than well-trained humans when the tasks are not known and are not part of the training data.
翻译:人工智能(AI)发展迅速,已展现出执行广泛认知任务的能力,包括语言处理、视觉识别和决策制定。这一进展部分归功于大语言模型(LLM),如GPT(生成式预训练Transformer)系列模型。这些模型能够表现出可被视为智能的行为。神经心理学领域的大多数学者认为,智能行为依赖于一系列总体性技能(即执行功能,EFs),这些功能依赖于额叶神经网络的正确运作,并已开发出系列测试来评估它们。本研究提出一个关键问题:大语言模型是否在学习过程中发展出类似人类的执行功能?我们使用经典的河内塔方法,评估了GPT的规划功能和工作记忆。此外,为避免解决方案存在于LLM训练数据中(数据泄露),我们引入了经典方法的新变体。初步结果显示,大语言模型在河内塔相关任务中生成接近最优解,能遵循任务约束,并展现出快速规划能力和高效工作记忆使用,表明其可能发展出执行功能。然而,当任务未知且未包含在训练数据中时,这些能力相当有限,且劣于训练有素的人类表现。