People make decisions differently in strategic interactions. Some update beliefs like a Bayesian; others exhibit biases like motivated reasoning. Although creators of large language models use simulated humans for safety evaluations and training, they often fail to cover this breadth of human behavior. We argue that cognitive science and economics provide a convenient tool for doing so, making use of mathematical models of human decision-making. We propose an approach that we call Equation-to-Behavior Prompting for guiding large language models to match cognitive models, and evaluate this approach on persuasion games based on legal decision-making. We find that large models can approximate equation-based specifications -- Bayesian updating, affine distortion, motivated updating, and Grether's $α$-$β$ model -- using prompting, but small models fail to do so. However, training small models with reinforcement learning to adhere to mathematical rules, Equation-to-Behavior RL, reduces belief error by 26.5% in out-of-distribution parameterizations. We show that these simulations can help create diverse training environments; training small models to consider different kinds of decision-makers improves average belief change by 2.5%--12% over Bayesian-only training, even when persuading GPT-5-mini. Our work could improve human simulations for training and evaluation in increasingly realistic settings, and could also enable novel research into more complicated mathematical models of human decision-making.
翻译:人们在战略互动中做出的决策各不相同。有些人像贝叶斯主义者一样更新信念;另一些人则表现出动机性推理等偏差。尽管大型语言模型的创建者使用模拟人类进行安全评估和训练,但他们往往未能涵盖人类行为的这种广度。我们认为,认知科学和经济学为此提供了便捷的工具,利用了人类决策的数学模型。我们提出了一种称为“方程到行为提示”的方法,用于引导大型语言模型匹配认知模型,并在基于法律决策的说服博弈中评估了该方法。我们发现,大型模型可以通过提示近似实现基于方程的规范——贝叶斯更新、仿射扭曲、动机更新以及Grether的$α$-$β$模型,但小型模型无法做到这一点。然而,通过强化学习训练小型模型使其遵循数学规则,即“方程到行为强化学习”,在分布外参数化中错误信念减少了26.5%。我们表明,这些模拟有助于创建多样化的训练环境;训练小型模型考虑不同类型的决策者,即使在与GPT-5-mini进行说服交互时,也使平均信念变化比仅用贝叶斯训练提高了2.5%–12%。我们的工作可以改进在日益逼真的环境中用于训练和评估的人类模拟,还可能为更复杂的人类决策数学模型研究开辟新途径。