Large language models (LLMs) such as ChatGPT and GPT-4 have recently demonstrated their remarkable abilities of communicating with human users. In this technical report, we take an initiative to investigate their capacities of playing text games, in which a player has to understand the environment and respond to situations by having dialogues with the game world. Our experiments show that ChatGPT performs competitively compared to all the existing systems but still exhibits a low level of intelligence. Precisely, ChatGPT can not construct the world model by playing the game or even reading the game manual; it may fail to leverage the world knowledge that it already has; it cannot infer the goal of each step as the game progresses. Our results open up new research questions at the intersection of artificial intelligence, machine learning, and natural language processing.
翻译:大型语言模型(如ChatGPT和GPT-4)近期展现出与人类用户进行对话的卓越能力。本技术报告率先探索了这类模型在文本游戏中的表现——玩家需通过理解游戏环境,以对话方式应对各类情景。实验表明,尽管ChatGPT的表现优于所有现有系统,但其智能水平仍显不足。具体而言:ChatGPT无法通过玩游戏或阅读游戏手册来构建世界模型;可能未能有效利用已具备的世界知识;也无法随游戏进程推断每一阶段的目标。我们的发现为人工智能、机器学习与自然语言处理交叉领域提出了新的研究课题。