Fair cooperative multi-agent reinforcement learning (MARL) teams that maximize an egalitarian welfare are exploitable: a single self-interested agent free-rides on the surplus that fair agents forgo to raise the worst-off, and the known remedy is a centralized need-based allocator. We show that a decentralized defense becomes possible once contention is graded: when a contested resource still delivers a fraction $1-c$, a worst-off cooperator that contests a free-rider strictly improves on yielding, so leverage exists for every $c < 1$. We introduce CAN, a permutation-equivariant cross-attention policy over agents' observed behaviour that infers how many free-riders are present and responds proportionally -- turn-taking when none, contesting just enough when some. Trained against an adversarial league, CAN keeps best-response exploitability near the centralized oracle ($ρ\approx 1.2\text{--}1.5$ vs. $ρ= N$ unprotected) at essentially no efficiency cost, whereas the fair-MARL learners (GGF, FEN, SOTO) each collapse to an exploitable or wasteful extreme. Giving those objectives CAN's identical adversarial training does not rescue them, so the objective -- not adversarial training alone -- is what makes hardening possible. Against a committed (non-adaptive) defector, every learned defense including ours provides deterrence rather than immunity, weakening as the leverage $(1-c)/2$ vanishes. Across further environments and team sizes the same principle sets the scope: robustness holds exactly as far as the game's contest leverage reaches, and we map that boundary rather than claim to remove it.
翻译:公平的合作多智能体强化学习(MARL)团队若最大化平等主义福利,则存在被利用的风险:一个自私的个体可通过搭便车行为坐享公平智能体为提高最差者福利而放弃的剩余,而已知的解决方法是采用集中式的按需分配器。我们证明,一旦资源竞争被分级,去中心化防御便成为可能:当被争夺资源仍能交付比例$1-c$时,与搭便车者竞争的最差劣势合作者比让步更能获益,因此对于每个$c < 1$都存在杠杆效应。我们提出CAN——一种基于智能体观测行为的排列等变交叉注意力策略,该策略能够推断出当前存在的搭便车者数量并做出相应响应:无搭车行为时采取轮换策略,存在部分搭车者时进行适度对抗。在与对抗联盟的训练中,CAN将最佳响应利用度保持在接近集中式基准的水平($ρ\approx 1.2\text{--}1.5$对比$ρ= N$的无保护状态),且几乎不损失效率;而公平MARL学习器(GGF、FEN、SOTO)则会陷入可被利用或效率低下的极端状态。即便对上述目标采用与CAN完全相同的对抗训练,也无法挽回其表现——因此,实现强化的关键在于目标本身而非对抗训练。面对固定(非自适应)叛逃者时,包括我们方法在内的所有学习型防御都仅能产生威慑而非免疫,其效果随杠杆系数$(1-c)/2$的消失而减弱。在更广泛的环境和团队规模中,相同原则划定了适用边界:鲁棒性恰好取决于博弈竞争杠杆所能触及的范围,我们描绘了这一边界而非声称能将其消除。