Categorical Distributional Reinforcement Learning (CDRL) has demonstrated superior sample efficiency in learning complex tasks compared to conventional Reinforcement Learning (RL) approaches. However, the practical application of CDRL is encumbered by challenging projection steps, detailed parameter tuning, and domain knowledge. This paper addresses these challenges by introducing a pioneering Continuous Distributional Model-Free RL algorithm tailored for continuous action spaces. The proposed algorithm simplifies the implementation of distributional RL, adopting an actor-critic architecture wherein the critic outputs a continuous probability distribution. Additionally, we propose an ensemble of multiple critics fused through a Kalman fusion mechanism to mitigate overestimation bias. Through a series of experiments, we validate that our proposed method is easy to train and serves as a sample-efficient solution for executing complex continuous-control tasks.
翻译:分类分布强化学习(CDRL)在学习复杂任务方面展现了比传统强化学习方法更高的样本效率。然而,CDRL的实际应用受到投影步骤复杂、参数调整繁琐以及领域知识依赖的制约。本文通过提出一种专为连续动作空间设计的前沿连续分布无模型强化学习算法来解决这些挑战。该算法简化了分布强化学习的实现,采用行动者-评论家架构,其中评论家输出连续概率分布。此外,我们提出了一种通过卡尔曼融合机制融合多个评论家的集成方法,以减轻过估计偏差。通过一系列实验,我们验证了所提方法易于训练,是执行复杂连续控制任务的一种样本高效解决方案。