We explore the intriguing possibility that theory of mind (ToM), or the uniquely human ability to impute unobservable mental states to others, might have spontaneously emerged in large language models (LLMs). We designed 40 false-belief tasks, considered a gold standard in testing ToM in humans, and administered them to several LLMs. Each task included a false-belief scenario, three closely matched true-belief controls, and the reversed versions of all four. Smaller and older models solved no tasks; GPT-3-davinci-003 (from November 2022) and ChatGPT-3.5-turbo (from March 2023) solved 20% of the tasks; ChatGPT-4 (from June 2023) solved 75% of the tasks, matching the performance of six-year-old children observed in past studies. These findings suggest the intriguing possibility that ToM, previously considered exclusive to humans, may have spontaneously emerged as a byproduct of LLMs' improving language skills.
翻译:我们探讨了一个引人入胜的可能性:心智理论(ToM),即人类将不可观察的心理状态归因于他人的独特能力,可能已在大语言模型(LLMs)中自发涌现。我们设计了40项错误信念任务(这些任务被视为测试人类心智理论的黄金标准),并对多个LLM进行了测试。每项任务包含一个错误信念场景、三个紧密匹配的真实信念控制场景,以及所有四个场景的反向版本。较小型和较老旧的模型未能解决任何任务;GPT-3-davinci-003(2022年11月版)和ChatGPT-3.5-turbo(2023年3月版)解决了20%的任务;ChatGPT-4(2023年6月版)解决了75%的任务,与过往研究中观察到的六岁儿童的表现相匹配。这些发现表明了一个引人入胜的可能性:此前被认为人类独有的心智理论,可能作为LLM语言能力提升的副产品而自发涌现。