Safe Reinforcement Learning (RL) plays an important role in applying RL algorithms to safety-critical real-world applications, addressing the trade-off between maximizing rewards and adhering to safety constraints. This work introduces a novel approach that combines RL with trajectory optimization to manage this trade-off effectively. Our approach embeds safety constraints within the action space of a modified Markov Decision Process (MDP). The RL agent produces a sequence of actions that are transformed into safe trajectories by a trajectory optimizer, thereby effectively ensuring safety and increasing training stability. This novel approach excels in its performance on challenging Safety Gym tasks, achieving significantly higher rewards and near-zero safety violations during inference. The method's real-world applicability is demonstrated through a safe and effective deployment in a real robot task of box-pushing around obstacles.
翻译:安全强化学习在将强化学习算法应用于安全关键型现实场景中发挥着重要作用,需权衡最大化奖励与遵守安全约束之间的矛盾。本文提出一种创新方法,通过将强化学习与轨迹优化相结合来有效管理这一权衡。该方法将安全约束嵌入到改进后的马尔可夫决策过程的动作空间中。强化学习智能体生成一系列动作,由轨迹优化器将其转化为安全轨迹,从而有效保证安全性并提升训练稳定性。该方法在具有挑战性的Safety Gym任务中表现出色,推理过程中实现了显著更高的奖励和近乎为零的安全违规。通过在实际机器人绕障推箱任务中的安全高效部署,验证了该方法在现实世界中的适用性。