Learning-based control algorithms have led to major advances in robotics at the cost of decreased safety guarantees. Recently, neural networks have also been used to characterize safety through the use of barrier functions for complex nonlinear systems. Learned barrier functions approximately encode and enforce a desired safety constraint through a value function, but do not provide any formal guarantees. In this paper, we propose a local dynamic programming (DP) based approach to "patch" an almost-safe learned barrier at potentially unsafe points in the state space. This algorithm, HJ-Patch, obtains a novel barrier that provides formal safety guarantees, yet retains the global structure of the learned barrier. Our local DP based reachability algorithm, HJ-Patch, updates the barrier function "minimally" at points that both (a) neighbor the barrier safety boundary and (b) do not satisfy the safety condition. We view this as a key step to bridging the gap between learning-based barrier functions and Hamilton-Jacobi reachability analysis, providing a framework for further integration of these approaches. We demonstrate that for well-trained barriers we reduce the computational load by 2 orders of magnitude with respect to standard DP-based reachability, and demonstrate scalability to a 6-dimensional system, which is at the limit of standard DP-based reachability.
翻译:基于学习的控制算法在机器人领域取得了重大进展,但代价是降低了安全性保证。近年来,神经网络也被用于通过障碍函数对复杂非线性系统进行安全性表征。学习得到的障碍函数通过值函数近似编码并实施期望的安全约束,但无法提供任何形式化保证。本文提出一种基于局部动态规划(DP)的方法,用于在状态空间中潜在不安全点处"修补"近似安全的已学习障碍函数。该算法HJ-Patch能够获得一个既提供形式化安全保证、又保留已学习障碍函数全局结构的新型障碍函数。我们提出的基于局部DP的可达性算法HJ-Patch,在同时满足以下两个条件的点处对障碍函数进行"最小化"更新:(a) 邻近障碍安全边界且(b) 不满足安全条件。这被视为弥合基于学习的障碍函数与Hamilton-Jacobi可达性分析之间差距的关键步骤,为这些方法的进一步融合提供了框架。实验表明,对于训练良好的障碍函数,相较于标准DP-based可达性计算,我们可将计算负载降低两个数量级,并展示了在六维系统(标准DP可达性方法的应用极限)上的可扩展性。