Multiple autonomous underwater vehicles (multi-AUV) can cooperatively accomplish tasks that a single AUV cannot complete. Recently, multi-agent reinforcement learning has been introduced to control of multi-AUV. However, designing efficient reward functions for various tasks of multi-AUV control is difficult or even impractical. Multi-agent generative adversarial imitation learning (MAGAIL) allows multi-AUV to learn from expert demonstration instead of pre-defined reward functions, but suffers from the deficiency of requiring optimal demonstrations and not surpassing provided expert demonstrations. This paper builds upon the MAGAIL algorithm by proposing multi-agent generative adversarial interactive self-imitation learning (MAGAISIL), which can facilitate AUVs to learn policies by gradually replacing the provided sub-optimal demonstrations with self-generated good trajectories selected by a human trainer. Our experimental results in a multi-AUV formation control and obstacle avoidance task on the Gazebo platform with AUV simulator of our lab show that AUVs trained via MAGAISIL can surpass the provided sub-optimal expert demonstrations and reach a performance close to or even better than MAGAIL with optimal demonstrations. Further results indicate that AUVs' policies trained via MAGAISIL can adapt to complex and different tasks as well as MAGAIL learning from optimal demonstrations.
翻译:多台自主水下航行器(multi-AUV)可协同完成单台AUV无法独立执行的任务。近年来,多智能体强化学习已被引入多AUV控制领域。然而,为多AUV控制的不同任务设计高效奖励函数具有较大难度甚至缺乏可行性。多智能体生成对抗模仿学习(MAGAIL)允许多AUV从专家示范中学习而无需预定义奖励函数,但其存在需要最优示范且无法超越所提供专家演示的缺陷。本文基于MAGAIL算法提出多智能体生成对抗交互式自我模仿学习(MAGAISIL),该方法通过由人类训练员筛选的自主生成优质轨迹逐步替代所提供的次优示范,从而促进AUV学习策略。在本实验室AUV仿真器与Gazebo平台上的多AUV编队控制与障碍规避任务实验中,经MAGAISIL训练的AUV不仅能够超越所提供的次优专家示范,其性能更接近甚至优于使用最优示范的MAGAIL。进一步结果表明,经MAGAISIL训练的AUV策略在适应复杂多样化任务方面,与使用最优示范学习的MAGAIL表现相当。