Theory of Mind (ToM) is the cognitive capability to perceive and ascribe mental states to oneself and others. Recent research has sparked a debate over whether large language models (LLMs) exhibit a form of ToM. However, existing ToM evaluations are hindered by challenges such as constrained scope, subjective judgment, and unintended contamination, yielding inadequate assessments. To address this gap, we introduce ToMBench with three key characteristics: a systematic evaluation framework encompassing 8 tasks and 31 abilities in social cognition, a multiple-choice question format to support automated and unbiased evaluation, and a build-from-scratch bilingual inventory to strictly avoid data leakage. Based on ToMBench, we conduct extensive experiments to evaluate the ToM performance of 10 popular LLMs across tasks and abilities. We find that even the most advanced LLMs like GPT-4 lag behind human performance by over 10% points, indicating that LLMs have not achieved a human-level theory of mind yet. Our aim with ToMBench is to enable an efficient and effective evaluation of LLMs' ToM capabilities, thereby facilitating the development of LLMs with inherent social intelligence.
翻译:心理理论(ToM)是一种感知并归因自我与他人心理状态的认知能力。近期研究引发了关于大型语言模型是否具备某种形式心理理论的讨论。然而,现有ToM评估面临范围受限、主观判断及意外污染等挑战,导致评估结果不够充分。为填补这一空白,我们提出ToMBench,其具备三个关键特征:包含社会认知领域8项任务与31种能力的系统化评估框架;支持自动化、无偏评估的多项选择题形式;以及从零构建的双语题库以严格避免数据泄露。基于ToMBench,我们开展了大规模实验,评估10个主流LLMs在不同任务与能力上的ToM表现。研究发现,即使最先进的模型(如GPT-4)在人类水平基准上仍落后超过10个百分点,表明LLMs尚未达到人类级别的心理理论能力。我们旨在通过ToMBench实现对LLM心理理论能力的高效准确评估,从而推动具备内在社会智能的语言模型发展。