Deep reinforcement learning (DRL) has achieved tremendous success in many complex decision-making tasks of autonomous systems with high-dimensional state and/or action spaces. However, the safety and stability still remain major concerns that hinder the applications of DRL to safety-critical autonomous systems. To address the concerns, we proposed the Phy-DRL: a physical deep reinforcement learning framework. The Phy-DRL is novel in two architectural designs: i) Lyapunov-like reward, and ii) residual control (i.e., integration of physics-model-based control and data-driven control). The concurrent physical reward and residual control empower the Phy-DRL the (mathematically) provable safety and stability guarantees. Through experiments on the inverted pendulum, we show that the Phy-DRL features guaranteed safety and stability and enhanced robustness, while offering remarkably accelerated training and enlarged reward.
翻译:深度强化学习(DRL)在高维状态和/或动作空间的自主系统复杂决策任务中取得了巨大成功。然而,安全性与稳定性仍然是阻碍DRL应用于安全关键型自主系统的主要问题。为解决这些问题,我们提出了Phy-DRL:一种物理深度强化学习框架。Phy-DRL在两种架构设计中具有新颖性:i)类李雅普诺夫奖励,以及ii)残差控制(即基于物理模型的控制与数据驱动控制的融合)。并行的物理奖励与残差控制赋予了Phy-DRL(数学上)可证明的安全性与稳定性保证。通过在倒立摆上的实验,我们展示了Phy-DRL具备可保证的安全性与稳定性以及增强的鲁棒性,同时实现了显著的训练加速和奖励提升。