Theory of Mind (ToM)$\unicode{x2014}$the ability to reason about the mental states of other people$\unicode{x2014}$is a key element of our social intelligence. Yet, despite their ever more impressive performance, large-scale neural language models still lack basic theory of mind capabilities out-of-the-box. We posit that simply scaling up models will not imbue them with theory of mind due to the inherently symbolic and implicit nature of the phenomenon, and instead investigate an alternative: can we design a decoding-time algorithm that enhances theory of mind of off-the-shelf neural language models without explicit supervision? We present SymbolicToM, a plug-and-play approach to reason about the belief states of multiple characters in reading comprehension tasks via explicit symbolic representation. More concretely, our approach tracks each entity's beliefs, their estimation of other entities' beliefs, and higher-order levels of reasoning, all through graphical representations, allowing for more precise and interpretable reasoning than previous approaches. Empirical results on the well-known ToMi benchmark (Le et al., 2019) demonstrate that SymbolicToM dramatically enhances off-the-shelf neural networks' theory of mind in a zero-shot setting while showing robust out-of-distribution performance compared to supervised baselines. Our work also reveals spurious patterns in existing theory of mind benchmarks, emphasizing the importance of out-of-distribution evaluation and methods that do not overfit a particular dataset.
翻译:心智理论(ToM)——即推理他人心理状态的能力——是我们社会智能的关键要素。然而,尽管大规模神经语言模型的表现日益惊人,它们仍然缺乏开箱即用的基础心智理论能力。我们认为,由于该现象固有的符号性和隐含性,单纯扩大模型规模无法赋予其心智理论,因此我们探索另一种方案:能否设计一种解码时间算法,在无需显式监督的情况下增强现成神经语言模型的心智理论?我们提出**SymbolicToM**,一种通过显式符号表征推理阅读理解任务中多角色信念状态的即插即用方法。更具体地说,我们的方法通过图形表征追踪每个实体的信念、它们对其他实体信念的估计以及更高阶的推理层次,从而比以往方法实现更精确且可解释的推理。在知名基准测试ToMi(Le等人,2019)上的实验结果表明,SymbolicToM在零样本设置下显著增强了现成神经网络的心智理论,同时与有监督基线相比展现了稳健的分布外性能。我们的工作还揭示了现有心智理论基准测试中的虚假模式,强调了分布外评估以及不针对特定数据集过拟合的方法的重要性。