Many safety-critical control systems must operate under latent uncertainty that sensors cannot directly resolve at decision time. Such uncertainty, arising from unknown physical properties, exogenous disturbances, or unobserved environment geometry, influences dynamics, task feasibility, and safety margins. Standard methods optimize expected performance and offer limited protection against rare but severe outcomes, while robust formulations treat uncertainty conservatively without exploiting its probabilistic structure. We consider partially observed dynamical systems whose dynamics, costs, and safety constraints depend on a latent parameter maintained as a belief distribution, and propose a risk-sensitive belief-space Model Predictive Path Integral (MPPI) control framework that plans under this belief while enforcing a Conditional Value-at-Risk (CVaR) constraint on a trajectory safety margin over the receding horizon. The resulting controller optimizes a risk-regularized performance objective while explicitly constraining the tail risk of safety violations induced by latent parameter variability. We establish three properties of the resulting risk-constrained controller: (1) the CVaR constraint implies a probabilistic safety guarantee, (2) the controller recovers the risk-neutral optimum as the risk weight in the objective tends to zero, and (3) a union-bound argument extends the per-horizon guarantee to cumulative safety over repeated solves. In physics-based simulations of a vision-guided dexterous stowing task in which a grasped object must be inserted into an occupied slot with pose uncertainty exceeding prescribed lateral clearance requirements, our method achieves 82% success with zero contact violations at high risk aversion, compared to 55% and 50% for a risk-neutral configuration and a chance-constrained baseline, both of which incur nonzero exterior contact forces.
翻译:许多安全关键控制系统必须在决策时传感器无法直接解析的潜变量不确定性下运行。这种不确定性源自未知物理属性、外部干扰或未观测的环境几何结构,会影响动力学、任务可行性及安全裕度。标准方法优化期望性能,但对罕见但严重的结果防护有限,而鲁棒性方法保守地处理不确定性,未利用其概率结构。我们考虑部分观测的动力学系统,其动态、成本和安全约束依赖于一个以信念分布形式维护的潜参数,并提出一种风险敏感的信念空间模型预测路径积分(MPPI)控制框架,该框架在该信念下进行规划,同时在对滚动时域内的轨迹安全裕度施加条件风险价值(CVaR)约束。由此产生的控制器优化风险正则化性能目标,同时明确约束由潜参数变异性引起的安全违规尾部风险。我们确立了所得风险约束控制器的三个性质:(1)CVaR约束蕴含概率安全性保证;(2)当目标中风险权重趋于零时,控制器恢复风险中性最优解;(3)通过联合界论证,将每个时域的保证扩展至重复求解的累积安全性。在视觉引导的灵巧存放任务的物理仿真中(被抓取物体需插入一个位姿不确定性超过预定侧向间隙要求的已占槽位),我们的方法在高风险规避下实现了82%的成功率且无接触违规,相比之下,风险中性配置和机会约束基线方法分别达到55%和50%,且两者均产生了非零外部接触力。