Recently, evolutionary reinforcement learning has obtained much attention in various domains. Maintaining a population of actors, evolutionary reinforcement learning utilises the collected experiences to improve the behaviour policy through efficient exploration. However, the poor scalability of genetic operators limits the efficiency of optimising high-dimensional neural networks. To address this issue, this paper proposes a novel cooperative coevolutionary reinforcement learning (CoERL) algorithm. Inspired by cooperative coevolution, CoERL periodically and adaptively decomposes the policy optimisation problem into multiple subproblems and evolves a population of neural networks for each of the subproblems. Instead of using genetic operators, CoERL directly searches for partial gradients to update the policy. Updating policy with partial gradients maintains consistency between the behaviour spaces of parents and offspring across generations. The experiences collected by the population are then used to improve the entire policy, which enhances the sampling efficiency. Experiments on six benchmark locomotion tasks demonstrate that CoERL outperforms seven state-of-the-art algorithms and baselines. Ablation study verifies the unique contribution of CoERL's core ingredients.
翻译:近年来,进化式强化学习在各个领域引起了广泛关注。通过维护一个智能体种群,进化式强化学习利用收集的经验通过高效探索来改进行为策略。然而,遗传算子扩展性差的问题限制了优化高维神经网络的效率。为解决这一问题,本文提出了一种新颖的合作协同进化强化学习(CoERL)算法。受合作协同进化启发,CoERL周期性地自适应地将策略优化问题分解为多个子问题,并为每个子问题进化一个神经网络种群。不同于使用遗传算子,CoERL直接搜索部分梯度来更新策略。基于部分梯度的策略更新保持了跨世代父子行为空间的一致性。随后,由种群收集的经验被用于改进整体策略,从而提升了采样效率。在六项基准运动控制任务上的实验表明,CoERL优于七种最先进算法及基线方法。消融研究验证了CoERL核心组件的独特贡献。