We introduce $λ$-Reachability, a scalable approach to Hamilton--Jacobi safety analysis for high-dimensional robotic systems. Unlike prior discounted formulations that rely on fixed one-step Bellman updates, $λ$-Reachability employs a stochastic multi-step estimator of the safety value, using a geometrically distributed rollout horizon together with a randomly absorbed terminal. Conceptually analogous to TD($λ$), $λ$-Reachability interpolates between local self-consistency updates and long-horizon max-over-trajectory safety targets via an interpretable horizon-control parameter. Unlike TD($λ$), where the terminal value is always incorporated in learning targets, the terminal safety value in $λ$-Reachability is only used at a probability controlled by parameter $δ$. We formally show that for $δ<1$, the update induces a contraction mapping that allows temporal-difference learning; as $λ\to 1$, the estimator recovers the undiscounted reachability objective. We apply $λ$-Reachability to high-dimensional safety learning problems with both simulated and real humanoid robots under balance and collision avoidance constraints. Experimental results demonstrate that $λ$-Reachability significantly improves both safe-set boundary classification and safety margin estimation compared to single-step temporal-difference baselines.
翻译:我们提出了$λ$-Reachability,一种面向高维机器人系统的哈密顿-雅可比安全性分析的可扩展方法。与依赖于固定一步贝尔曼更新的传统折扣公式不同,$λ$-Reachability采用安全值的随机多步估计器,该方法利用几何分布的展开视界以及随机吸收的终止状态。在概念上类似于TD($λ$),$λ$-Reachability通过一个可解释的视界控制参数,在局部自洽性更新与长期轨迹最大化安全目标之间进行插值。与TD($λ$)不同(后者总是将终止值纳入学习目标),$λ$-Reachability中终止安全值仅以参数$δ$控制的概率被使用。我们严格证明了:当$δ<1$时,该更新诱导出一个压缩映射,从而支持时序差分学习;当$λ\to 1$时,估计器恢复无折扣的可达性目标。我们将$λ$-Reachability应用于高维安全学习问题,在平衡与避碰约束条件下,分别使用仿真和真实人形机器人进行实验。实验结果表明,与单步时序差分基线相比,$λ$-Reachability显著提升了安全集边界分类与安全裕度估计的性能。