Reputation, the aggregation of peer assessments diffused through social networks, is a pivotal mechanism for promoting cooperation in social dilemmas ubiquitous to distributed multi-agent systems comprising agents with limited perception and cognitive capabilities. Exploring efficient reputation systems, comprising reputation assessment rules and reputation-based policies, is a long-standing challenge. Previous work assumes predefined reputation assessment rules or models reputation as an intrinsic reward to learn policies, compromising the methods' ability for generalization and adaptation. To address this, we propose a distributed multi-agent reinforcement learning method $\textbf{COOPER}$ ($\textbf{COOP}$eration with $\textbf{E}$mergent $\textbf{R}$eputation), which jointly learns reputation assessment rules and reputation-based policies entirely from environment rewards. Notably, leveraging the underlying mechanisms of reputation, we deliberately design the constituent modules of $\textbf{COOPER}$ and the data flows among them, overcoming the latency and noise in the feedback signal, caused by the deep entanglement between reputation and policy. Experiments on the donation game and the coin game in grid world environments demonstrate that $\textbf{COOPER}$ effectively adapts to various existing reputation systems and co-players. Furthermore, we observe the co-emergence of reputation norms and cooperation in self-play settings. These results hold robustly across diverse social network topologies, underscoring the generalizability and efficacy of our approach.
翻译:声誉,即通过社交网络扩散的同伴评估的聚合,是促进分布式多智能体系统中(由感知和认知能力受限的智能体组成)社会困境中合作的关键机制。探索高效的声誉系统(包括声誉评估规则和基于声誉的策略)是一个长期挑战。以往研究假定预设的声誉评估规则或将声誉建模为内在奖励以学习策略,这损害了方法的泛化与适应能力。为解决此问题,我们提出一种分布式多智能体强化学习方法$\textbf{COOPER}$($\textbf{COOP}$eration with $\textbf{E}$mergent $\textbf{R}$eputation),该方法完全从环境奖励中联合学习声誉评估规则和基于声誉的策略。值得注意的是,我们利用声誉的底层机制,精心设计了$\textbf{COOPER}$的组成模块及其间数据流,克服了声誉与策略深度纠缠导致的反馈信号延迟与噪声。在网格世界环境中的捐赠游戏和硬币游戏实验表明,$\textbf{COOPER}$能有效适应各种现有声誉系统及合作对手。此外,我们在自对弈设置中观察到声誉规范与合作的共同涌现。这些结果在不同社交网络拓扑下保持稳健,凸显了我们方法的泛化性和有效性。