Theory of mind (ToM), or the ability to impute unobservable mental states to others, is central to human social interactions, communication, empathy, self-consciousness, and morality. We administer classic false-belief tasks, widely used to test ToM in humans, to several language models, without any examples or pre-training. Our results show that models published before 2022 show virtually no ability to solve ToM tasks. Yet, the January 2022 version of GPT-3 (davinci-002) solved 70% of ToM tasks, a performance comparable with that of seven-year-old children. Moreover, its November 2022 version (davinci-003), solved 93% of ToM tasks, a performance comparable with that of nine-year-old children. These findings suggest that ToM-like ability (thus far considered to be uniquely human) may have spontaneously emerged as a byproduct of language models' improving language skills.
翻译:心智理论(Theory of Mind, ToM),即对他人不可观测心理状态进行归因的能力,是人类社交互动、沟通、共情、自我意识及道德认知的核心。我们采用经典错误信念任务(广泛用于人类ToM测试)对多个语言模型进行测试,未使用任何示例或预训练。结果表明,2022年前发布的模型几乎不具备解决ToM任务的能力。然而,2022年1月版本的GPT-3(davinci-002)解决了70%的ToM任务,其表现与七岁儿童相当。更值得注意的是,其2022年11月版本(davinci-003)解决了93%的ToM任务,表现与九岁儿童相当。这些发现表明,类似心智理论的能力(迄今被认为人类独有)可能作为语言模型语言技能提升的副产品而自发涌现。