We study general-sum, multi-player stochastic games with transferable utility, motivated by settings where agents can use side payments to make cooperation individually rational. Building on the Harsanyi--Shapley (HS) value for normal-form games, we introduce two HS-based value notions for stochastic games: HS-S, defined by aggregating dynamic coalition-versus-complement threat powers, and Coco-S, defined as fixed points of a statewise HS Bellman operator. We extend HS-style axioms to the stochastic setting and show that HS-S is the unique mapping satisfying them. We prove that HS-S and Coco-S coincide in all two-player stochastic games, but can disagree when $n>2$, via an explicit three-player counterexample. We prove existence and uniqueness of Coco-S fixed points for all two-player games and for three-player two-state games via topological degree theory, and provide an axiomatic characterization of Coco-S through a new \emph{Markov Consistency} axiom that distinguishes it from HS-S. Finally, we give sampling-based estimators with finite-sample guarantees and empirically compare the induced values, policies, and side payments on multi-player grid-game benchmarks.


翻译:我们研究具有可转移效用的多人、一般和随机博弈,其动机源于智能体可以通过辅助支付使合作变得个人理性的场景。基于规范形式博弈的Harsanyi-Shapley(HS)价值,我们引入了两种用于随机博弈的基于HS的价值概念:HS-S(通过聚合动态联盟与互补威胁权能定义)和Coco-S(定义为状态wise HS贝尔曼算子的不动点)。我们将HS风格的公理扩展到随机环境,并证明HS-S是唯一满足这些公理的映射。我们证明在所有两人随机博弈中HS-S与Coco-S是等价的,但当$n>2$时,通过一个显式的三人反例表明二者可能不一致。利用拓扑度理论,我们证明了所有两人博弈及三人两状态博弈中Coco-S不动点的存在性与唯一性,并通过一项新的\emph{马尔可夫一致性}公理(该公理将Coco-S与HS-S区分开来)给出了Coco-S的公理化刻画。最后,我们给出了具有有限样本保证的基于采样的估计量,并在多人网格博弈基准上实证比较了诱导出的价值、策略及辅助支付。

0
下载
关闭预览

相关内容

基于多智能体强化学习的博弈综述
专知会员服务
53+阅读 · 2024年11月23日
多智能体博弈中的分布式学习: 原理与算法
专知会员服务
54+阅读 · 2024年6月13日
多智能体博弈学习研究进展
专知会员服务
91+阅读 · 2024年5月5日
博弈论应用《互补战场上的多场战斗对抗》
专知会员服务
27+阅读 · 2024年1月30日
多智能体学习中合作的综述
专知会员服务
75+阅读 · 2023年12月12日
【2023新书】合作博弈论的计算方面,170页pdf
专知会员服务
72+阅读 · 2023年6月29日
面向智能博弈的决策Transformer方法综述
专知会员服务
201+阅读 · 2023年4月14日
多智能体博弈、学习与控制
专知会员服务
129+阅读 · 2023年1月18日
面向多智能体博弈对抗的对手建模框架
专知
18+阅读 · 2022年9月28日
半监督多任务学习:Semisupervised Multitask Learning
我爱读PAMI
18+阅读 · 2018年4月29日
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Arxiv
0+阅读 · 5月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
8+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
6+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
13+阅读 · 7月31日
相关VIP内容
基于多智能体强化学习的博弈综述
专知会员服务
53+阅读 · 2024年11月23日
多智能体博弈中的分布式学习: 原理与算法
专知会员服务
54+阅读 · 2024年6月13日
多智能体博弈学习研究进展
专知会员服务
91+阅读 · 2024年5月5日
博弈论应用《互补战场上的多场战斗对抗》
专知会员服务
27+阅读 · 2024年1月30日
多智能体学习中合作的综述
专知会员服务
75+阅读 · 2023年12月12日
【2023新书】合作博弈论的计算方面,170页pdf
专知会员服务
72+阅读 · 2023年6月29日
面向智能博弈的决策Transformer方法综述
专知会员服务
201+阅读 · 2023年4月14日
多智能体博弈、学习与控制
专知会员服务
129+阅读 · 2023年1月18日
相关基金
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员