Reinforcement learning has gained significant traction in the field of robotic navigation. However, a persistent challenge is its sample inefficiency, primarily due to the inherent complexities of encouraging exploration. During training, the mobile agent must explore as much as possible to efficiently learn optimal behaviors. We introduce Ada-NAV, a novel adaptive trajectory length scheme designed to enhance the training sample efficiency of reinforcement learning algorithms in robotic navigation tasks. Unlike traditional approaches that treat trajectory length as a fixed hyperparameter, Ada-NAV dynamically adjusts it based on the entropy of the underlying navigation policy. We empirically validate the efficacy of AdaNAV using two popular policy gradient methods: REINFORCE and Proximal Policy Optimization (PPO). We demonstrate through both simulated and real-world robotic experiments that Ada-NAV outperforms conventional methods that employ constant or randomly sampled trajectory lengths. Specifically, for a fixed sample budget, Ada-NAV achieves an 18% increase in navigation success rate, a 20-38% reduction in navigation path length, and a 9.32% decrease in elevation costs. Furthermore, we showcase the versatility of Ada-NAV by integrating it with the Clearpath Husky robot, illustrating its applicability in complex, outdoor environments.
翻译:强化学习在机器人导航领域取得了显著进展。然而,其样本效率低下的问题始终存在,主要源于鼓励探索的内在复杂性。在训练过程中,移动智能体必须尽可能充分地探索,以高效学习最优行为。我们提出Ada-NAV——一种新颖的自适应轨迹长度方案,旨在提升机器人导航任务中强化学习算法的训练样本效率。与将轨迹长度视为固定超参数的传统方法不同,Ada-NAV根据底层导航策略的熵动态调整轨迹长度。我们采用两种主流的策略梯度方法——REINFORCE和近端策略优化(PPO)——实证验证了Ada-NAV的有效性。通过仿真及真实机器人实验证明,Ada-NAV在性能上优于采用恒定或随机采样轨迹长度的传统方法。具体而言,在固定样本预算下,Ada-NAV使导航成功率提升18%,导航路径长度缩短20%-38%,海拔成本降低9.32%。此外,我们通过将Ada-NAV集成至Clearpath Husky机器人,展示了其在复杂户外环境中的适用性。