Learning long-horizon tasks such as navigation has presented difficult challenges for successfully applying reinforcement learning to robotics. From another perspective, under known environments, sampling-based planning can robustly find collision-free paths in environments without learning. In this work, we propose Control Transformer that models return-conditioned sequences from low-level policies guided by a sampling-based Probabilistic Roadmap (PRM) planner. We demonstrate that our framework can solve long-horizon navigation tasks using only local information. We evaluate our approach on partially-observed maze navigation with MuJoCo robots, including Ant, Point, and Humanoid. We show that Control Transformer can successfully navigate through mazes and transfer to unknown environments. Additionally, we apply our method to a differential drive robot (Turtlebot3) and show zero-shot sim2real transfer under noisy observations.
翻译:学习诸如导航等长时域任务为成功将强化学习应用于机器人技术带来了巨大挑战。另一方面,在已知环境下,基于采样的规划方法能够在不依赖学习的情况下鲁棒地找到无碰撞路径。本文提出控制变换器,该模型通过由基于采样的概率路图(PRM)规划器引导的低层策略建模返回条件序列。我们证明该框架仅利用局部信息即可解决长时域导航任务。我们在基于MuJoCo机器人的部分可观测迷宫导航场景中评估该方法,包括Ant、Point和Humanoid。实验表明控制变换器能够成功穿越迷宫并迁移至未知环境。此外,我们将该方法应用于差速驱动机器人(Turtlebot3),并在噪声观测下实现了零样本仿真到现实的迁移。