Maneuver decision-making is the core of unmanned combat aerial vehicle for autonomous air combat. To solve this problem, we propose an automatic curriculum reinforcement learning method, which enables agents to learn effective decisions in air combat from scratch. The range of initial states are used for distinguishing curricula of different difficulty levels, thereby maneuver decision is divided into a series of sub-tasks from easy to difficult, and test results are used to change sub-tasks. As sub-tasks change, agents gradually learn to complete a series of sub-tasks from easy to difficult, enabling them to make effective maneuvering decisions to cope with various states without the need to spend effort designing reward functions. The ablation studied show that the automatic curriculum learning proposed in this article is an essential component for training through reinforcement learning, namely, agents cannot complete effective decisions without curriculum learning. Simulation experiments show that, after training, agents are able to make effective decisions given different states, including tracking, attacking and escaping, which are both rational and interpretable.
翻译:机动决策是无人作战飞行器实现自主空战的核心。针对该问题,我们提出一种自动课程强化学习方法,使智能体能够从零开始在空战中学习有效决策。利用初始状态范围区分不同难度等级的课程,从而将机动决策分解为一系列从易到难的子任务,并通过测试结果实现子任务切换。随着子任务的变化,智能体逐步学习完成从易到难的系列子任务,使其无需耗费精力设计奖励函数即可应对多种状态进行有效机动决策。消融研究表明,本文提出的自动课程学习是通过强化学习训练的关键组件——若缺乏课程学习机制,智能体无法完成有效决策。仿真实验表明,经过训练的智能体能够在包括追踪、攻击和逃逸等不同状态下作出合理且可解释的有效决策。