Strategic dialogue requires agents to execute distinct dialogue acts, for which belief estimation is essential. While prior work often estimates beliefs accurately, it lacks a principled mechanism to use those beliefs during generation. We bridge this gap by first formalizing two core acts Adversarial and Alignment, and by operationalizing them via probabilistic constraints on what an agent may generate. We instantiate this idea in BEDA, a framework that consists of the world set, the belief estimator for belief estimation, and the conditional generator that selects acts and realizes utterances consistent with the inferred beliefs. Across three settings, Conditional Keeper Burglar (CKBG, adversarial), Mutual Friends (MF, cooperative), and CaSiNo (negotiation), BEDA consistently outperforms strong baselines: on CKBG it improves success rate by at least 5.0 points across backbones and by 20.6 points with GPT-4.1-nano; on Mutual Friends it achieves an average improvement of 9.3 points; and on CaSiNo it achieves the optimal deal relative to all baselines. These results indicate that casting belief estimation as constraints provides a simple, general mechanism for reliable strategic dialogue.


翻译:策略性对话要求智能体执行不同的对话行为,信念估计对此至关重要。虽然先前的研究通常能准确估计信念,但缺乏在生成过程中利用这些信念的机制。我们通过以下方式弥补这一空白:首先形式化两种核心行为——对抗与对齐,并通过智能体生成内容的概率约束将其操作化。我们在BEDA框架中实现了这一思想,该框架包含世界集合、用于信念估计的信念估计器,以及根据推断信念选择行为并生成一致话语的条件生成器。在条件守护者-盗贼(CKBG,对抗性)、共同好友(MF,合作性)和CaSiNo(协商性)三种设定中,BEDA始终优于强基线模型:在CKBG任务中,其在各骨干模型上成功率至少提升5.0个百分点,使用GPT-4.1-nano时提升20.6个百分点;在共同好友任务中平均提升9.3个百分点;在CaSiNo任务中达成了相对于所有基线的最优协议。这些结果表明,将信念估计转化为约束条件为可靠的策略性对话提供了一种简单通用的机制。

0
下载
关闭预览

相关内容

《面向人机协作的扩展型信念-愿望-意图模型》最新111页
《人工智能辅助决策中信任的时间演化​​》225页
专知会员服务
25+阅读 · 2025年5月12日
因果决策综述
专知会员服务
51+阅读 · 2025年3月1日
《基于信念的决策建模计算框架》141页
专知会员服务
71+阅读 · 2024年4月27日
《OODA 和 CECA:决策框架分析》
专知会员服务
116+阅读 · 2023年11月8日
探索(Exploration)还是利用(Exploitation)?强化学习如何tradeoff?
深度强化学习实验室
13+阅读 · 2020年8月23日
强化学习的两大话题之一,仍有极大探索空间
AI科技评论
22+阅读 · 2020年8月22日
赛尔原创 | 对话系统评价方法综述
哈工大SCIR
11+阅读 · 2017年11月13日
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2012年12月31日
国家自然科学基金
97+阅读 · 2009年12月31日
国家自然科学基金
36+阅读 · 2008年12月31日
VIP会员
最新内容
美空军新型反无人机部队初探
专知会员服务
3+阅读 · 今天5:45
《防空交战流程的概率建模研究》
专知会员服务
6+阅读 · 今天5:04
ICML 2026 教程 | 数值优化理论还重要吗?
专知会员服务
4+阅读 · 7月26日
ICM 2026 | 陶哲轩:人工智能时代的数学
专知会员服务
7+阅读 · 7月26日
《反无人机交战场景下的战斗归零研究》
专知会员服务
7+阅读 · 7月26日
博士论文 | 用代码结构感知方法推进代码大模型
相关基金
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2012年12月31日
国家自然科学基金
97+阅读 · 2009年12月31日
国家自然科学基金
36+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员