Fair cooperative multi-agent reinforcement learning (MARL) teams that maximize an egalitarian welfare are exploitable: a single self-interested agent free-rides on the surplus that fair agents forgo to raise the worst-off, and the known remedy is a centralized need-based allocator. We show that a decentralized defense becomes possible once contention is graded: when a contested resource still delivers a fraction $1-c$, a worst-off cooperator that contests a free-rider strictly improves on yielding, so leverage exists for every $c < 1$. We introduce CAN, a permutation-equivariant cross-attention policy over agents' observed behaviour that infers how many free-riders are present and responds proportionally -- turn-taking when none, contesting just enough when some. Trained against an adversarial league, CAN keeps best-response exploitability near the centralized oracle ($ρ\approx 1.2\text{--}1.5$ vs. $ρ= N$ unprotected) at essentially no efficiency cost, whereas the fair-MARL learners (GGF, FEN, SOTO) each collapse to an exploitable or wasteful extreme. Giving those objectives CAN's identical adversarial training does not rescue them, so the objective -- not adversarial training alone -- is what makes hardening possible. Against a committed (non-adaptive) defector, every learned defense including ours provides deterrence rather than immunity, weakening as the leverage $(1-c)/2$ vanishes. Across further environments and team sizes the same principle sets the scope: robustness holds exactly as far as the game's contest leverage reaches, and we map that boundary rather than claim to remove it.


翻译:公平的合作多智能体强化学习(MARL)团队若最大化平等主义福利,则存在被利用的风险:一个自私的个体可通过搭便车行为坐享公平智能体为提高最差者福利而放弃的剩余,而已知的解决方法是采用集中式的按需分配器。我们证明,一旦资源竞争被分级,去中心化防御便成为可能:当被争夺资源仍能交付比例$1-c$时,与搭便车者竞争的最差劣势合作者比让步更能获益,因此对于每个$c < 1$都存在杠杆效应。我们提出CAN——一种基于智能体观测行为的排列等变交叉注意力策略,该策略能够推断出当前存在的搭便车者数量并做出相应响应:无搭车行为时采取轮换策略,存在部分搭车者时进行适度对抗。在与对抗联盟的训练中,CAN将最佳响应利用度保持在接近集中式基准的水平($ρ\approx 1.2\text{--}1.5$对比$ρ= N$的无保护状态),且几乎不损失效率;而公平MARL学习器(GGF、FEN、SOTO)则会陷入可被利用或效率低下的极端状态。即便对上述目标采用与CAN完全相同的对抗训练,也无法挽回其表现——因此,实现强化的关键在于目标本身而非对抗训练。面对固定(非自适应)叛逃者时,包括我们方法在内的所有学习型防御都仅能产生威慑而非免疫,其效果随杠杆系数$(1-c)/2$的消失而减弱。在更广泛的环境和团队规模中,相同原则划定了适用边界:鲁棒性恰好取决于博弈竞争杠杆所能触及的范围,我们描绘了这一边界而非声称能将其消除。

0
下载
关闭预览

相关内容

多智能体强化学习中的稳健且高效的通信
专知会员服务
26+阅读 · 2025年11月17日
《分布式多智能体强化学习策略的可解释性研究》
专知会员服务
30+阅读 · 2025年11月17日
开放环境下的协作多智能体强化学习进展综述
专知会员服务
35+阅读 · 2025年1月19日
多智能体学习中合作的综述
专知会员服务
75+阅读 · 2023年12月12日
「基于通信的多智能体强化学习」 进展综述
面向多智能体博弈对抗的对手建模框架
专知
18+阅读 · 2022年9月28日
【综述】多智能体强化学习算法理论研究
深度强化学习实验室
16+阅读 · 2020年9月9日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
38+阅读 · 2020年3月15日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
8+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
7+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
14+阅读 · 7月31日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员