Reinforcement learning (RL) excels in applications such as video games and robotics, but ensuring safety and stability remains challenging when using RL to control real-world systems where using model-free algorithms suffering from low sample efficiency might be prohibitive. This paper first provides safety and stability definitions for the RL system, and then introduces a Neural ordinary differential equations-based Lyapunov-Barrier Actor-Critic (NLBAC) framework that leverages Neural Ordinary Differential Equations (NODEs) to approximate system dynamics and integrates the Control Barrier Function (CBF) and Control Lyapunov Function (CLF) frameworks with the actor-critic method to assist in maintaining the safety and stability for the system. Within this framework, we employ the augmented Lagrangian method to update the RL-based controller parameters. Additionally, we introduce an extra backup controller in situations where CBF constraints for safety and the CLF constraint for stability cannot be satisfied simultaneously. Simulation results demonstrate that the framework leads the system to approach the desired state and allows fewer violations of safety constraints with better sample efficiency compared to other methods.
翻译:强化学习在视频游戏和机器人等应用中表现出色,但在使用强化学习控制真实世界系统时,确保安全性和稳定性仍面临挑战——尤其在采用样本效率低下的无模型算法可能不可行的情况下。本文首先给出了强化学习系统的安全性与稳定性定义,随后提出了一种基于神经常微分方程的Lyapunov-屏障Actor-Critic(NLBAC)框架。该框架利用神经常微分方程(NODEs)逼近系统动态,并将控制屏障函数(CBF)与控制Lyapunov函数(CLF)框架与actor-critic方法相结合,以辅助维持系统的安全性和稳定性。在此框架中,我们采用增广拉格朗日方法更新基于强化学习的控制器参数。此外,当安全性CBF约束与稳定性CLF约束无法同时满足时,我们引入了一个额外的备份控制器。仿真结果表明,与其他方法相比,该框架能够引导系统趋近期望状态,并在降低安全约束违反次数的同时实现更优的样本效率。