People make decisions differently in strategic interactions. Some update beliefs like a Bayesian; others exhibit biases like motivated reasoning. Although creators of large language models use simulated humans for safety evaluations and training, they often fail to cover this breadth of human behavior. We argue that cognitive science and economics provide a convenient tool for doing so, making use of mathematical models of human decision-making. We propose an approach that we call Equation-to-Behavior Prompting for guiding large language models to match cognitive models, and evaluate this approach on persuasion games based on legal decision-making. We find that large models can approximate equation-based specifications -- Bayesian updating, affine distortion, motivated updating, and Grether's $α$-$β$ model -- using prompting, but small models fail to do so. However, training small models with reinforcement learning to adhere to mathematical rules, Equation-to-Behavior RL, reduces belief error by 26.5% in out-of-distribution parameterizations. We show that these simulations can help create diverse training environments; training small models to consider different kinds of decision-makers improves average belief change by 2.5%--12% over Bayesian-only training, even when persuading GPT-5-mini. Our work could improve human simulations for training and evaluation in increasingly realistic settings, and could also enable novel research into more complicated mathematical models of human decision-making.


翻译:人们在战略互动中做出的决策各不相同。有些人像贝叶斯主义者一样更新信念;另一些人则表现出动机性推理等偏差。尽管大型语言模型的创建者使用模拟人类进行安全评估和训练,但他们往往未能涵盖人类行为的这种广度。我们认为,认知科学和经济学为此提供了便捷的工具,利用了人类决策的数学模型。我们提出了一种称为“方程到行为提示”的方法,用于引导大型语言模型匹配认知模型,并在基于法律决策的说服博弈中评估了该方法。我们发现,大型模型可以通过提示近似实现基于方程的规范——贝叶斯更新、仿射扭曲、动机更新以及Grether的$α$-$β$模型,但小型模型无法做到这一点。然而,通过强化学习训练小型模型使其遵循数学规则,即“方程到行为强化学习”,在分布外参数化中错误信念减少了26.5%。我们表明,这些模拟有助于创建多样化的训练环境;训练小型模型考虑不同类型的决策者,即使在与GPT-5-mini进行说服交互时,也使平均信念变化比仅用贝叶斯训练提高了2.5%–12%。我们的工作可以改进在日益逼真的环境中用于训练和评估的人类模拟,还可能为更复杂的人类决策数学模型研究开辟新途径。

0
下载
关闭预览

相关内容

【NTU博士论文】让语言模型成为更类人的学习者
专知会员服务
23+阅读 · 2025年9月23日
【NTU博士论文】让语言模型更接近人类学习者
专知会员服务
18+阅读 · 2025年5月3日
博弈论与大语言模型的结合:系统性综述
专知会员服务
60+阅读 · 2025年2月14日
《大型语言模型情感认知》最新进展
专知会员服务
43+阅读 · 2024年10月3日
《兵棋推演与大型语言模型: 方法、应用和稳健性》
专知会员服务
38+阅读 · 2024年7月19日
【博士论文】语言模型与人类偏好对齐,148页pdf
专知会员服务
32+阅读 · 2024年4月21日
《人类与机器:大型语言模型与兵棋推演》
专知会员服务
90+阅读 · 2024年3月27日
通过大型语言模型的双重使用加速认知战
专知会员服务
67+阅读 · 2024年1月17日
大型语言模型被称为太空部队的 “游戏规则改变器”
专知会员服务
34+阅读 · 2023年12月15日
「知识增强预训练语言模型」最新研究综述
专知
18+阅读 · 2022年11月18日
绝对干货!NLP预训练模型:从transformer到albert
新智元
14+阅读 · 2019年11月10日
进一步改进GPT和BERT:使用Transformer的语言模型
机器之心
16+阅读 · 2019年5月1日
自然语言处理中的语言模型预训练方法
PaperWeekly
14+阅读 · 2018年10月21日
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
51+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
43+阅读 · 2012年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
9+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
9+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
14+阅读 · 7月31日
相关VIP内容
【NTU博士论文】让语言模型成为更类人的学习者
专知会员服务
23+阅读 · 2025年9月23日
【NTU博士论文】让语言模型更接近人类学习者
专知会员服务
18+阅读 · 2025年5月3日
博弈论与大语言模型的结合:系统性综述
专知会员服务
60+阅读 · 2025年2月14日
《大型语言模型情感认知》最新进展
专知会员服务
43+阅读 · 2024年10月3日
《兵棋推演与大型语言模型: 方法、应用和稳健性》
专知会员服务
38+阅读 · 2024年7月19日
【博士论文】语言模型与人类偏好对齐,148页pdf
专知会员服务
32+阅读 · 2024年4月21日
《人类与机器:大型语言模型与兵棋推演》
专知会员服务
90+阅读 · 2024年3月27日
通过大型语言模型的双重使用加速认知战
专知会员服务
67+阅读 · 2024年1月17日
大型语言模型被称为太空部队的 “游戏规则改变器”
专知会员服务
34+阅读 · 2023年12月15日
相关基金
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
51+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
43+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员