Collision-free, goal-directed navigation in environments containing unknown static and dynamic obstacles is still a great challenge, especially when manual tuning of navigation policies or costly motion prediction needs to be avoided. In this paper, we therefore propose a subgoal-driven hierarchical navigation architecture that is trained with deep reinforcement learning and decouples obstacle avoidance and motor control. In particular, we separate the navigation task into the prediction of the next subgoal position for avoiding collisions while moving toward the final target position, and the prediction of the robot's velocity controls. By relying on 2D lidar, our method learns to avoid obstacles while still achieving goal-directed behavior as well as to generate low-level velocity control commands to reach the subgoals. In our architecture, we apply the attention mechanism on the robot's 2D lidar readings and compute the importance of lidar scan segments for avoiding collisions. As we show in simulated and real-world experiments with a Turtlebot robot, our proposed method leads to smooth and safe trajectories among humans and significantly outperforms a state-of-the-art approach in terms of success rate. A supplemental video describing our approach is available online.
翻译:在包含未知静态和动态障碍物的环境中实现无碰撞、目标导向的导航仍是一项重大挑战,尤其是需要避免手动调整导航策略或进行昂贵的运动预测时。为此,本文提出一种基于子目标驱动、采用深度强化学习训练的分层导航架构,将障碍物规避与运动控制解耦。具体而言,我们将导航任务分解为两个子任务:预测下一个子目标位置(在向最终目标移动时避免碰撞)与预测机器人的速度控制指令。该方法仅依赖2D激光雷达,即可学习在保持目标导向行为的同时规避障碍物,并生成底层速度控制指令以抵达子目标。在我们的架构中,注意力机制被应用于机器人的2D激光雷达读数,用于计算激光雷达扫描片段对避免碰撞的重要性。通过Turtlebot机器人的仿真与真实世界实验,我们提出的方法能生成平滑且安全的人类周边运动轨迹,并在成功率上显著优于现有最优方法。补充视频可在线上获取。