Reinforcement learning (RL) is a machine learning approach that trains agents to maximize cumulative rewards through interactions with environments. The integration of RL with deep learning has recently resulted in impressive achievements in a wide range of challenging tasks, including board games, arcade games, and robot control. Despite these successes, there remain several crucial challenges, including brittle convergence properties caused by sensitive hyperparameters, difficulties in temporal credit assignment with long time horizons and sparse rewards, a lack of diverse exploration, especially in continuous search space scenarios, difficulties in credit assignment in multi-agent reinforcement learning, and conflicting objectives for rewards. Evolutionary computation (EC), which maintains a population of learning agents, has demonstrated promising performance in addressing these limitations. This article presents a comprehensive survey of state-of-the-art methods for integrating EC into RL, referred to as evolutionary reinforcement learning (EvoRL). We categorize EvoRL methods according to key research fields in RL, including hyperparameter optimization, policy search, exploration, reward shaping, meta-RL, and multi-objective RL. We then discuss future research directions in terms of efficient methods, benchmarks, and scalable platforms. This survey serves as a resource for researchers and practitioners interested in the field of EvoRL, highlighting the important challenges and opportunities for future research. With the help of this survey, researchers and practitioners can develop more efficient methods and tailored benchmarks for EvoRL, further advancing this promising cross-disciplinary research field.
翻译:强化学习(RL)是一种通过智能体与环境交互来最大化累积奖励的机器学习方法。深度学习与强化学习的结合近期在包括棋类游戏、街机游戏和机器人控制等广泛挑战性任务中取得了令人瞩目的成果。尽管取得了这些成功,但依然存在若干关键挑战,包括超参数敏感性导致的脆性收敛特性、长时域与稀疏奖励条件下的时序信用分配困难、探索多样性不足(尤其在连续搜索空间场景中)、多智能体强化学习中的信用分配问题以及奖励目标的冲突性。进化计算(EC)通过维护学习智能体种群的方式,在解决上述局限性方面展现出良好潜力。本文全面综述了将进化计算融入强化学习(称为进化强化学习)的最先进方法。我们根据强化学习的关键研究领域对进化强化学习方法进行分类,包括超参数优化、策略搜索、探索、奖励塑形、元强化学习和多目标强化学习。随后探讨了高效方法、基准测试和可扩展平台等未来研究方向。本综述为关注进化强化学习领域的研究者与从业者提供了重要参考,揭示了该领域的研究挑战与未来机遇。借助本综述,研究者与从业者可开发更高效的进化强化学习方法与定制化基准测试,进一步推动这一前景广阔的跨学科研究领域的发展。