Reinforcement learning (RL) has demonstrated impressive performance in various areas such as video games and robotics. However, ensuring safety and stability, which are two critical properties from a control perspective, remains a significant challenge when using RL to control real-world systems. In this paper, we first provide definitions of safety and stability for the RL system, and then combine the control barrier function (CBF) and control Lyapunov function (CLF) methods with the actor-critic method in RL to propose a Barrier-Lyapunov Actor-Critic (BLAC) framework which helps maintain the aforementioned safety and stability for the system. In this framework, CBF constraints for safety and CLF constraint for stability are constructed based on the data sampled from the replay buffer, and the augmented Lagrangian method is used to update the parameters of the RL-based controller. Furthermore, an additional backup controller is introduced in case the RL-based controller cannot provide valid control signals when safety and stability constraints cannot be satisfied simultaneously. Simulation results show that this framework yields a controller that can help the system approach the desired state and cause fewer violations of safety constraints compared to baseline algorithms.
翻译:强化学习(RL)在电子游戏和机器人等多个领域展现了令人瞩目的性能。然而,在用RL控制真实世界系统时,如何确保安全性和稳定性(这两个从控制视角出发的关键特性)仍然是一项重大挑战。本文首先为RL系统定义了安全性和稳定性的概念,随后将控制障碍函数(CBF)和控制李雅普诺夫函数(CLF)方法与RL中的演员-评论家方法相结合,提出了障碍-李雅普诺夫演员-评论家(BLAC)框架,该框架有助于维持系统的上述安全性和稳定性。在该框架中,基于从回放缓冲区采样的数据构建了用于安全性的CBF约束和用于稳定性的CLF约束,并采用增广拉格朗日方法更新基于RL控制器的参数。此外,当安全性和稳定性约束无法同时满足导致基于RL控制器无法提供有效控制信号时,引入了额外的备用控制器。仿真结果表明,与基准算法相比,该框架生成的控制器既能帮助系统趋近期望状态,又能减少违反安全约束的次数。