Reinforcement learning (RL) has achieved promising results on most robotic control tasks. Safety of learning-based controllers is an essential notion of ensuring the effectiveness of the controllers. Current methods adopt whole consistency constraints during the training, thus resulting in inefficient exploration in the early stage. In this paper, we propose an algorithm named Constrained Policy Optimization with Extra Safety Budget (ESB-CPO) to strike a balance between the exploration efficiency and the constraints satisfaction. In the early stage, our method loosens the practical constraints of unsafe transitions (adding extra safety budget) with the aid of a new metric we propose. With the training process, the constraints in our optimization problem become tighter. Meanwhile, theoretical analysis and practical experiments demonstrate that our method gradually meets the cost limit's demand in the final training stage. When evaluated on Safety-Gym and Bullet-Safety-Gym benchmarks, our method has shown its advantages over baseline algorithms in terms of safety and optimality. Remarkably, our method gains remarkable performance improvement under the same cost limit compared with baselines.
翻译:强化学习(RL)在大多数机器人控制任务中取得了有希望的结果。基于学习的控制器的安全性是确保控制器有效性的关键概念。当前方法在训练过程中采用全程一致性约束,导致早期探索效率低下。本文提出一种名为“带额外安全预算的约束策略优化(ESB-CPO)”的算法,以平衡探索效率与约束满足。在早期阶段,我们的方法借助提出的新度量指标,放宽了不安全转移的实际约束(增加额外安全预算)。随着训练过程进行,优化问题中的约束逐渐收紧。同时,理论分析和实践实验表明,我们的方法在最终训练阶段逐步满足成本限制的要求。在Safety-Gym和Bullet-Safety-Gym基准测试中评估时,我们的方法在安全性和最优性方面展示了相对于基线算法的优势。值得注意的是,与基线方法相比,在相同成本限制下,我们的方法获得了显著的性能提升。