Safety assurance is uncompromisable for safety-critical environments with the presence of drastic model uncertainties (e.g., distributional shift), especially with humans in the loop. However, incorporating uncertainty in safe learning will naturally lead to a bi-level problem, where at the lower level the (worst-case) safety constraint is evaluated within the uncertainty ambiguity set. In this paper, we present a tractable distributionally safe reinforcement learning framework to enforce safety under a distributional shift measured by a Wasserstein metric. To improve the tractability, we first use duality theory to transform the lower-level optimization from infinite-dimensional probability space where distributional shift is measured, to a finite-dimensional parametric space. Moreover, by differentiable convex programming, the bi-level safe learning problem is further reduced to a single-level one with two sequential computationally efficient modules: a convex quadratic program to guarantee safety followed by a projected gradient ascent to simultaneously find the worst-case uncertainty. This end-to-end differentiable framework with safety constraints, to the best of our knowledge, is the first tractable single-level solution to address distributional safety. We test our approach on first and second-order systems with varying complexities and compare our results with the uncertainty-agnostic policies, where our approach demonstrates a significant improvement on safety guarantees.
翻译:在存在剧烈模型不确定性(如分布偏移)的安全关键环境中(尤其涉及人机交互时),安全性保障是不可妥协的。然而,在安全学习中纳入不确定性会自然导致双层优化问题:下层需在不确定性模糊集内评估(最坏情况下的)安全约束。本文提出一种易于处理的分布鲁棒安全强化学习框架,以在Wasserstein度量衡量的分布偏移下保障安全性。为提升可解性,我们首先利用对偶理论将下层优化从度量分布偏移的无限维概率空间转换至有限维参数空间。进一步地,通过可微凸规划,双层安全学习问题被简化为单层问题,包含两个顺序耦合的高效计算模块:保障安全性的凸二次规划模块,以及同步寻找最坏情况不确定性的投影梯度上升模块。据我们所知,这一具有安全约束的端到端可微框架是首个处理分布安全性问题的可解单层解决方案。我们在具有不同复杂度的二阶系统上测试了该方法,并将结果与不确定性无关策略进行对比,实验表明我们的方法在安全保证方面具有显著提升。