Hierarchical reinforcement learning (HRL) has led to remarkable achievements in diverse fields. However, existing HRL algorithms still cannot be applied to real-world navigation tasks. These tasks require an agent to perform safety-aware behaviors and interact with surrounding objects in dynamic environments. In addition, an agent in these tasks should perform consistent and structured exploration as they are long-horizon and have complex structures with diverse objects and task-specific rules. Designing HRL agents that can handle these challenges in real-world navigation tasks is an open problem. In this paper, we propose imagination-augmented HRL (IAHRL), a new and general navigation algorithm that allows an agent to learn safe and interactive behaviors in real-world navigation tasks. Our key idea is to train a hierarchical agent in which a high-level policy infers interactions by interpreting behaviors imagined with low-level policies. Specifically, the high-level policy is designed with a permutation-invariant attention mechanism to determine which low-level policy generates the most interactive behavior, and the low-level policies are implemented with an optimization-based behavior planner to generate safe and structured behaviors following task-specific rules. To evaluate our algorithm, we introduce five complex urban driving tasks, which are among the most challenging real-world navigation tasks. The experimental results indicate that our hierarchical agent performs safety-aware behaviors and properly interacts with surrounding vehicles, achieving higher success rates and lower average episode steps than baselines in urban driving tasks.
翻译:分层强化学习(HRL)已在多个领域取得显著成就。然而,现有HRL算法仍无法应用于真实世界的导航任务。这类任务要求智能体在动态环境中执行安全感知行为并与周围物体进行交互。此外,由于这些任务具有长时域特性且包含复杂结构(涉及多种物体及任务特定规则),智能体应执行一致且结构化的探索。设计能够应对真实导航任务中这些挑战的HRL智能体仍是一个开放性问题。本文提出想象力增强分层强化学习(IAHRL)——一种新颖且通用的导航算法,使智能体能够在真实导航任务中学习安全交互行为。我们的核心思想是训练一个分层智能体,其中高层策略通过解释低层策略想象的行为来推断交互。具体而言,高层策略采用置换不变注意力机制,用于判定哪个低层策略能产生最具交互性的行为;低层策略则通过基于优化的行为规划器实现,以生成遵循任务特定规则的安全结构化行为。为评估算法性能,我们引入了五个复杂的城市驾驶任务——这些属于最具挑战性的真实导航任务之一。实验结果表明,我们的分层智能体能够执行安全感知行为并与周围车辆进行恰当交互,在城市驾驶任务中实现了比基线方法更高的成功率和更低的平均回合步数。