Large language models (LLMs) are increasingly tasked with strategic decision-making under incomplete information, such as in negotiation and policymaking. While LLMs can excel at many such tasks, they also fail in ways that are poorly understood. We shed light on these failures by uncovering two fundamental gaps in the internal mechanisms underlying the decision-making of LLMs in incomplete-information games, supported by experiments with open-weight models Llama 3.1, Qwen3, and gpt-oss. First, an observation-belief gap: LLMs encode internal beliefs about latent game states that are substantially more accurate than their own verbal reports, yet these beliefs are brittle. In particular, the belief accuracy degrades with multi-hop reasoning, exhibits primacy and recency biases, and drifts away from Bayesian coherence over extended interactions. Second, a belief-action gap: The implicit conversion of internal beliefs into actions is weaker than that of the beliefs externalized in the prompt, yet neither belief-conditioning consistently achieves higher game payoffs. These results show how analyzing LLMs' internal processes can expose systematic vulnerabilities that warrant caution before deploying LLMs in strategic domains without robust guardrails.
翻译:大型语言模型(LLMs)越来越多地承担起在信息不完全条件下的战略决策任务,例如谈判和政策制定。尽管LLMs在许多此类任务中表现出色,但它们也会以我们尚未充分理解的方式失败。我们通过揭示LLMs在不完全信息博弈中决策背后内部机制的两个根本性缺陷来阐明这些失败,并利用开源权重模型Llama 3.1、Qwen3和gpt-oss进行的实验加以佐证。首先,存在观察-信念缺口:LLMs关于潜在博弈状态的内部编码信念,比其自身的口头报告准确得多,然而这些信念是脆弱的。具体而言,信念的准确性会随着多跳推理而下降,表现出首因效应和近因效应偏差,并在长时间交互中偏离贝叶斯一致性。其次,存在信念-行动缺口:将内部信念隐式转化为行动的能力弱于将提示中外显化信念转化为行动的能力,然而,这两种信念调节方式均未能持续获得更高的博弈收益。这些结果表明,分析LLMs的内部过程能够揭示系统性脆弱点,这提醒我们在没有稳健护栏的情况下,将LLMs部署到战略领域时应保持谨慎。