In social dilemmas self-interested learning agents face the choice between the societal benefit of cooperation and the immediate reward of defection. Significant evidence exists on the benefits of assortment mechanisms such as partner selection for the emergence of cooperation, but this is largely available through agent-based simulations. In this paper, we provide an analytical solution to the problem, studying the policy-gradient dynamics in a multi-agent environment with partner selection. We show how partner selection changes the opponent distribution and hence the reward landscape, and prove this promotes cooperation under simple rules known from the literature. In particular, we find that population variance is a necessary condition for cooperation to emerge. Using a two-dimensional Wiener process, we extend the dynamics to capture the stochastic effects of partner selection and the resulting opponent distribution. We derive a sufficient condition for the population to be cooperation-promoting and prove the existence of a stationary distribution. Simulations confirm that the stochastic model accurately captures the policy-gradient dynamics and clarifies how the learning rate affects the emergence of cooperation.
翻译:在具有自利学习主体的社会困境中,主体面临合作的社会效益与背叛的即时奖励之间的抉择。大量证据表明,伙伴选择等分类机制有利于合作涌现,但这些证据主要来自基于主体建模的仿真。本文为该问题提供了解析解,研究带有伙伴选择机制的多主体环境中的策略梯度动力学。我们展示了伙伴选择如何改变对手分布,进而重塑奖励景观,并证明在文献已知的简单规则下,这一机制能够促进合作。特别地,我们发现群体方差是合作涌现的必要条件。通过引入二维维纳过程,我们将动力学扩展至捕捉伙伴选择及其所导致的对手分布中的随机效应。我们推导出群体具有促进合作功能的充分条件,并证明平稳分布的存在性。仿真结果验证了该随机模型能够准确刻画策略梯度动力学,并阐明了学习率如何影响合作涌现。