The False Belief Test (FBT) has been the main method for assessing Theory of Mind (ToM) and related socio-cognitive competencies. For Large Language Models (LLMs), the reliability and explanatory potential of this test have remained limited due to issues like data contamination, insufficient model details, and inconsistent controls. We address these issues by testing 17 open-weight models on a balanced set of 192 FBT variants (Trott et al., 2023) using Bayesian Logistic regression to identify how model size and post-training affect socio-cognitive competence. We find that scaling model size benefits performance, but not strictly. A cross-over effect reveals that explicating propositional attitudes (X thinks) fundamentally alters response patterns. Instruction tuning partially mitigates this effect, but further reasoning-oriented fine-tuning amplifies it. In a case study analysing social reasoning ability throughout OLMo 2 training, we show that this cross-over effect emerges during pre-training, suggesting that models acquire stereotypical response patterns tied to mental-state vocabulary that can outweigh other scenario semantics. Finally, vector steering allows us to isolate a think vector as the causal driver of observed FBT behaviour.


翻译:虚假信念测试(False Belief Test, FBT)一直是评估心理理论(Theory of Mind, ToM)及相关社会认知能力的主要方法。对于大型语言模型(LLMs),由于数据污染、模型细节不足以及控制条件不一致等问题,该测试的可靠性和解释潜力仍受到限制。我们通过在一组平衡的192个FBT变体(Trott 等人,2023)上测试17个开源权重模型来解决这些问题,并采用贝叶斯逻辑回归来识别模型规模和后训练对社会认知能力的影响。我们发现,扩大模型规模有利于性能提升,但并非绝对如此。一种交叉效应表明,明确表达命题态度(如“X认为”)从根本上改变了响应模式。指令微调部分缓解了这种效应,但进一步以推理为导向的微调则会放大它。在分析OLMo 2训练过程中社会推理能力的案例研究中,我们展示了这种交叉效应在预训练阶段便已出现,表明模型获得了与心理状态词汇相关的刻板响应模式,这些模式可能会压倒其他场景语义。最后,向量引导使我们能够将“思维向量”隔离为观察到FBT行为的因果驱动因素。

0
下载
关闭预览

相关内容

基于大语言模型智能体的社会认知模拟
专知会员服务
20+阅读 · 2月22日
大语言模型的智能体化推理
专知会员服务
36+阅读 · 1月21日
评估大语言模型在科学发现中的作用
专知会员服务
19+阅读 · 2025年12月19日
大型语言模型的规模效应局限
专知会员服务
14+阅读 · 2025年11月18日
基于大型语言模型的人机系统综述
专知会员服务
26+阅读 · 2025年5月12日
大型语言模型中的人格综述
专知会员服务
42+阅读 · 2024年6月30日
「大型语言模型推理」综述
专知会员服务
96+阅读 · 2022年12月24日
「知识增强预训练语言模型」最新研究综述
专知
18+阅读 · 2022年11月18日
稀疏大模型简述:从MoE、Sparse Attention到GLaM
夕小瑶的卖萌屋
14+阅读 · 2022年3月22日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Arxiv
26+阅读 · 2024年2月9日
Arxiv
21+阅读 · 2023年7月12日
Arxiv
25+阅读 · 2023年6月23日
Arxiv
12+阅读 · 2023年5月31日
VIP会员
最新内容
反制无人机:乌克兰提供的五点启示
专知会员服务
1+阅读 · 今天14:52
《各指挥层级均亟需红队能力》报告
专知会员服务
2+阅读 · 今天14:46
《远程传感对战机架次生成影响的仿真研究》80页
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
4+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
7+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
9+阅读 · 9月22日
战争不仅需要机器人:人类仍不可或缺
专知会员服务
5+阅读 · 9月21日
《描绘美国防部创新基础设施的未来蓝图》100页
专知会员服务
10+阅读 · 9月21日
相关VIP内容
基于大语言模型智能体的社会认知模拟
专知会员服务
20+阅读 · 2月22日
大语言模型的智能体化推理
专知会员服务
36+阅读 · 1月21日
评估大语言模型在科学发现中的作用
专知会员服务
19+阅读 · 2025年12月19日
大型语言模型的规模效应局限
专知会员服务
14+阅读 · 2025年11月18日
基于大型语言模型的人机系统综述
专知会员服务
26+阅读 · 2025年5月12日
大型语言模型中的人格综述
专知会员服务
42+阅读 · 2024年6月30日
「大型语言模型推理」综述
专知会员服务
96+阅读 · 2022年12月24日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员