The False Belief Test (FBT) has been the main method for assessing Theory of Mind (ToM) and related socio-cognitive competencies. For Large Language Models (LLMs), the reliability and explanatory potential of this test have remained limited due to issues like data contamination, insufficient model details, and inconsistent controls. We address these issues by testing 17 open-weight models on a balanced set of 192 FBT variants (Trott et al., 2023) using Bayesian Logistic regression to identify how model size and post-training affect socio-cognitive competence. We find that scaling model size benefits performance, but not strictly. A cross-over effect reveals that explicating propositional attitudes (X thinks) fundamentally alters response patterns. Instruction tuning partially mitigates this effect, but further reasoning-oriented fine-tuning amplifies it. In a case study analysing social reasoning ability throughout OLMo 2 training, we show that this cross-over effect emerges during pre-training, suggesting that models acquire stereotypical response patterns tied to mental-state vocabulary that can outweigh other scenario semantics. Finally, vector steering allows us to isolate a think vector as the causal driver of observed FBT behaviour.
翻译:虚假信念测试(False Belief Test, FBT)一直是评估心理理论(Theory of Mind, ToM)及相关社会认知能力的主要方法。对于大型语言模型(LLMs),由于数据污染、模型细节不足以及控制条件不一致等问题,该测试的可靠性和解释潜力仍受到限制。我们通过在一组平衡的192个FBT变体(Trott 等人,2023)上测试17个开源权重模型来解决这些问题,并采用贝叶斯逻辑回归来识别模型规模和后训练对社会认知能力的影响。我们发现,扩大模型规模有利于性能提升,但并非绝对如此。一种交叉效应表明,明确表达命题态度(如“X认为”)从根本上改变了响应模式。指令微调部分缓解了这种效应,但进一步以推理为导向的微调则会放大它。在分析OLMo 2训练过程中社会推理能力的案例研究中,我们展示了这种交叉效应在预训练阶段便已出现,表明模型获得了与心理状态词汇相关的刻板响应模式,这些模式可能会压倒其他场景语义。最后,向量引导使我们能够将“思维向量”隔离为观察到FBT行为的因果驱动因素。