We study general-sum, multi-player stochastic games with transferable utility, motivated by settings where agents can use side payments to make cooperation individually rational. Building on the Harsanyi--Shapley (HS) value for normal-form games, we introduce two HS-based value notions for stochastic games: HS-S, defined by aggregating dynamic coalition-versus-complement threat powers, and Coco-S, defined as fixed points of a statewise HS Bellman operator. We extend HS-style axioms to the stochastic setting and show that HS-S is the unique mapping satisfying them. We prove that HS-S and Coco-S coincide in all two-player stochastic games, but can disagree when $n>2$, via an explicit three-player counterexample. We prove existence and uniqueness of Coco-S fixed points for all two-player games and for three-player two-state games via topological degree theory, and provide an axiomatic characterization of Coco-S through a new \emph{Markov Consistency} axiom that distinguishes it from HS-S. Finally, we give sampling-based estimators with finite-sample guarantees and empirically compare the induced values, policies, and side payments on multi-player grid-game benchmarks.
翻译:我们研究具有可转移效用的多人、一般和随机博弈,其动机源于智能体可以通过辅助支付使合作变得个人理性的场景。基于规范形式博弈的Harsanyi-Shapley(HS)价值,我们引入了两种用于随机博弈的基于HS的价值概念:HS-S(通过聚合动态联盟与互补威胁权能定义)和Coco-S(定义为状态wise HS贝尔曼算子的不动点)。我们将HS风格的公理扩展到随机环境,并证明HS-S是唯一满足这些公理的映射。我们证明在所有两人随机博弈中HS-S与Coco-S是等价的,但当$n>2$时,通过一个显式的三人反例表明二者可能不一致。利用拓扑度理论,我们证明了所有两人博弈及三人两状态博弈中Coco-S不动点的存在性与唯一性,并通过一项新的\emph{马尔可夫一致性}公理(该公理将Coco-S与HS-S区分开来)给出了Coco-S的公理化刻画。最后,我们给出了具有有限样本保证的基于采样的估计量,并在多人网格博弈基准上实证比较了诱导出的价值、策略及辅助支付。