Decentralized multi-agent reinforcement learning (MARL) algorithms have become popular in the literature since it allows heterogeneous agents to have their own reward functions as opposed to canonical multi-agent Markov Decision Process (MDP) settings which assume common reward functions over all agents. In this work, we follow the existing work on collaborative MARL where agents in a connected time varying network can exchange information among each other in order to reach a consensus. We introduce vulnerabilities in the consensus updates of existing MARL algorithms where agents can deviate from their usual consensus update, who we term as adversarial agents. We then proceed to provide an algorithm that allows non-adversarial agents to reach a consensus in the presence of adversaries under a constrained setting.
翻译:分散式多智能体强化学习(MARL)算法在文献中已日趋流行,它允许异构智能体拥有各自的奖励函数,这与传统多智能体马尔可夫决策过程(MDP)设置不同——后者假设所有智能体共享同一奖励函数。本研究延续了协作式MARL的现有工作,其中智能体在连通的时变网络中可相互交换信息以达成共识。我们发现现有MARL算法的共识更新环节存在漏洞:部分智能体(称为对抗性智能体)可能偏离常规的共识更新规则。据此,我们提出一种算法,使得在受限环境下,非对抗性智能体即便面对对抗性智能体也能达成共识。