Safety in goal directed Reinforcement Learning (RL) settings has typically been handled through constraints over trajectories and have demonstrated good performance in primarily short horizon tasks. In this paper, we are specifically interested in the problem of solving temporally extended decision making problems such as robots cleaning different areas in a house while avoiding slippery and unsafe areas (e.g., stairs) and retaining enough charge to move to a charging dock; in the presence of complex safety constraints. Our key contribution is a (safety) Constrained Search with Hierarchical Reinforcement Learning (CoSHRL) mechanism that combines an upper level constrained search agent (which computes a reward maximizing policy from a given start to a far away goal state while satisfying cost constraints) with a low-level goal conditioned RL agent (which estimates cost and reward values to move between nearby states). A major advantage of CoSHRL is that it can handle constraints on the cost value distribution (e.g., on Conditional Value at Risk, CVaR) and can adjust to flexible constraint thresholds without retraining. We perform extensive experiments with different types of safety constraints to demonstrate the utility of our approach over leading approaches in constrained and hierarchical RL.
翻译:目标导向强化学习中的安全性通常通过轨迹约束实现,并在短时域任务中展现出良好性能。本文重点研究存在复杂安全约束时,解决时间维度延展决策问题的方法,例如机器人需在规避湿滑区域和危险区域(如楼梯)的同时保持足够电量返回充电基座,完成房屋不同区域的清洁任务。我们提出的核心贡献是(安全)约束搜索与分层强化学习结合机制(CoSHRL),该机制将上层约束搜索智能体(在满足成本约束条件下,从给定起点到远端目标状态计算奖励最大化策略)与底层目标条件强化学习智能体(估计相邻状态间的转移成本与奖励值)进行整合。CoSHRL的主要优势在于能够处理成本值分布上的约束(如条件风险价值上的约束),且无需重新训练即可适应灵活变化的约束阈值。我们通过涵盖不同类型的复杂安全约束实验,验证了本方法相较于约束强化学习与分层强化学习领域主流方法的优越性。