As safety is of paramount importance in robotics, reinforcement learning that reflects safety, called safe RL, has been studied extensively. In safe RL, we aim to find a policy which maximizes the desired return while satisfying the defined safety constraints. There are various types of constraints, among which constraints on conditional value at risk (CVaR) effectively lower the probability of failures caused by high costs since CVaR is a conditional expectation obtained above a certain percentile. In this paper, we propose a trust region-based safe RL method with CVaR constraints, called TRC. We first derive the upper bound on CVaR and then approximate the upper bound in a differentiable form in a trust region. Using this approximation, a subproblem to get policy gradients is formulated, and policies are trained by iteratively solving the subproblem. TRC is evaluated through safe navigation tasks in simulations with various robots and a sim-to-real environment with a Jackal robot from Clearpath. Compared to other safe RL methods, the performance is improved by 1.93 times while the constraints are satisfied in all experiments.
翻译:由于安全性在机器人领域至关重要,反映安全性的强化学习(称为安全强化学习)已得到广泛研究。在安全强化学习中,我们旨在寻找一种既能最大化期望回报又能满足既定安全约束的策略。约束类型多样,其中条件风险价值(CVaR)约束能有效降低由高成本导致的故障概率,因为CVaR是超过某一百分位数得到的条件期望。本文提出一种基于置信域且具有CVaR约束的安全强化学习方法,称为TRC。我们首先推导出CVaR的上界,然后在置信域内以可微形式逼近该上界。利用这一近似,构建了获取策略梯度的子问题,并通过迭代求解该子问题来训练策略。我们通过多种机器人的仿真安全导航任务以及搭载Clearpath公司Jackal机器人的仿真-真实环境对TRC进行评估。与其他安全强化学习方法相比,该方法在所有实验中均满足约束条件的同时,性能提升了1.93倍。