We study the problem of multi-agent coordination in unpredictable and partially-observable environments with untrustworthy external commands. The commands are actions suggested to the robots, and are untrustworthy in that their performance guarantees, if any, are unknown. Such commands may be generated by human operators or machine learning algorithms and, although untrustworthy, can often increase the robots' performance in complex multi-robot tasks. We are motivated by complex multi-robot tasks such as target tracking, environmental mapping, and area monitoring. Such tasks are often modeled as submodular maximization problems due to the information overlap among the robots. We provide an algorithm, Meta Bandit Sequential Greedy (MetaBSG), which enjoys performance guarantees even when the external commands are arbitrarily bad. MetaBSG leverages a meta-algorithm to learn whether the robots should follow the commands or a recently developed submodular coordination algorithm, Bandit Sequential Greedy (BSG) [1], which has performance guarantees even in unpredictable and partially-observable environments. Particularly, MetaBSG asymptotically can achieve the better performance out of the commands and the BSG algorithm, quantifying its suboptimality against the optimal time-varying multi-robot actions in hindsight. Thus, MetaBSG can be interpreted as robustifying the untrustworthy commands. We validate our algorithm in simulated scenarios of multi-target tracking.
翻译:我们研究了在不可预测且部分可观测的环境中,利用不可靠外部指令实现多智能体协作的问题。指令是向机器人建议的动作,其不可靠性体现在其性能保证(若存在)未知。此类指令可能由人类操作员或机器学习算法生成,尽管不可靠,但在复杂的多机器人任务中通常能提升机器人性能。我们的研究受多目标跟踪、环境地图绘制和区域监控等复杂多机器人任务驱动。由于机器人之间存在信息重叠,这些任务通常被建模为子模最大化问题。我们提出了一种名为元强盗序列贪婪(Meta Bandit Sequential Greedy, MetaBSG)的算法,即使外部指令任意恶劣,该算法仍能保持性能保证。MetaBSG利用元算法学习机器人应遵循指令,还是采用近期开发的子模协调算法——强盗序列贪婪(Bandit Sequential Greedy, BSG)[1],该算法在不可预测且部分可观测的环境中同样具有性能保证。特别地,MetaBSG能渐进地实现指令与BSG算法中较优者的性能,并量化其相对于事后最优时变多机器人动作的次优性。因此,MetaBSG可被视为对不可靠指令的鲁棒化处理。我们通过多目标跟踪的模拟场景验证了该算法。