Reputation, the aggregation of peer assessments diffused through social networks, is a pivotal mechanism for promoting cooperation in social dilemmas ubiquitous to distributed multi-agent systems comprising agents with limited perception and cognitive capabilities. Exploring efficient reputation systems, comprising reputation assessment rules and reputation-based policies, is a long-standing challenge. Previous work assumes predefined reputation assessment rules or models reputation as an intrinsic reward to learn policies, compromising the methods' ability for generalization and adaptation. To address this, we propose a distributed multi-agent reinforcement learning method $\textbf{COOPER}$ ($\textbf{COOP}$eration with $\textbf{E}$mergent $\textbf{R}$eputation), which jointly learns reputation assessment rules and reputation-based policies entirely from environment rewards. Notably, leveraging the underlying mechanisms of reputation, we deliberately design the constituent modules of $\textbf{COOPER}$ and the data flows among them, overcoming the latency and noise in the feedback signal, caused by the deep entanglement between reputation and policy. Experiments on the donation game and the coin game in grid world environments demonstrate that $\textbf{COOPER}$ effectively adapts to various existing reputation systems and co-players. Furthermore, we observe the co-emergence of reputation norms and cooperation in self-play settings. These results hold robustly across diverse social network topologies, underscoring the generalizability and efficacy of our approach.


翻译:声誉,即通过社交网络扩散的同伴评估的聚合,是促进分布式多智能体系统中(由感知和认知能力受限的智能体组成)社会困境中合作的关键机制。探索高效的声誉系统(包括声誉评估规则和基于声誉的策略)是一个长期挑战。以往研究假定预设的声誉评估规则或将声誉建模为内在奖励以学习策略,这损害了方法的泛化与适应能力。为解决此问题,我们提出一种分布式多智能体强化学习方法$\textbf{COOPER}$($\textbf{COOP}$eration with $\textbf{E}$mergent $\textbf{R}$eputation),该方法完全从环境奖励中联合学习声誉评估规则和基于声誉的策略。值得注意的是,我们利用声誉的底层机制,精心设计了$\textbf{COOPER}$的组成模块及其间数据流,克服了声誉与策略深度纠缠导致的反馈信号延迟与噪声。在网格世界环境中的捐赠游戏和硬币游戏实验表明,$\textbf{COOPER}$能有效适应各种现有声誉系统及合作对手。此外,我们在自对弈设置中观察到声誉规范与合作的共同涌现。这些结果在不同社交网络拓扑下保持稳健,凸显了我们方法的泛化性和有效性。

0
下载
关闭预览

相关内容

多智能体协作机制
专知会员服务
25+阅读 · 4月25日
面向关系建模的合作多智能体深度强化学习综述
专知会员服务
42+阅读 · 2025年4月18日
《多智能体合作强化学习中的通信》139页
专知会员服务
47+阅读 · 2025年2月17日
《多智能体强化学习的深度合作策略》最新154页博士论文
专知会员服务
64+阅读 · 2024年11月18日
基于学习机制的多智能体强化学习综述
专知会员服务
64+阅读 · 2024年4月16日
多智能体学习中合作的综述
专知会员服务
75+阅读 · 2023年12月12日
基于多智能体强化学习的协同目标分配
专知会员服务
142+阅读 · 2023年9月5日
「基于通信的多智能体强化学习」 进展综述
面向多智能体博弈对抗的对手建模框架
专知
18+阅读 · 2022年9月28日
浅谈群体智能——新一代AI的重要方向
中国科学院自动化研究所
44+阅读 · 2019年10月16日
基于车路协同的群体智能协同
智能交通技术
10+阅读 · 2019年1月23日
群体智能:新一代人工智能的重要方向
走向智能论坛
12+阅读 · 2017年8月16日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
VIP会员
最新内容
非对称防御中的自组织临界性:俄乌战争
专知会员服务
8+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
8+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
11+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
12+阅读 · 8月8日
相关VIP内容
多智能体协作机制
专知会员服务
25+阅读 · 4月25日
面向关系建模的合作多智能体深度强化学习综述
专知会员服务
42+阅读 · 2025年4月18日
《多智能体合作强化学习中的通信》139页
专知会员服务
47+阅读 · 2025年2月17日
《多智能体强化学习的深度合作策略》最新154页博士论文
专知会员服务
64+阅读 · 2024年11月18日
基于学习机制的多智能体强化学习综述
专知会员服务
64+阅读 · 2024年4月16日
多智能体学习中合作的综述
专知会员服务
75+阅读 · 2023年12月12日
基于多智能体强化学习的协同目标分配
专知会员服务
142+阅读 · 2023年9月5日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员