Autonomous robots operating in complex environments face the critical challenge of identifying and utilizing environmental cover for covert navigation to minimize exposure to potential threats. We propose EnCoMP, an enhanced navigation framework that integrates offline reinforcement learning and our novel Adaptive Threat-Aware Visibility Estimation (ATAVE) algorithm to enable robots to navigate covertly and efficiently in diverse outdoor settings. ATAVE is a dynamic probabilistic threat modeling technique that we designed to continuously assess and mitigate potential threats in real-time, enhancing the robot's ability to navigate covertly by adapting to evolving environmental and threat conditions. Moreover, our approach generates high-fidelity multi-map representations, including cover maps, potential threat maps, height maps, and goal maps from LiDAR point clouds, providing a comprehensive understanding of the environment. These multi-maps offer detailed environmental insights, helping in strategic navigation decisions. The goal map encodes the relative distance and direction to the target location, guiding the robot's navigation. We train a Conservative Q-Learning (CQL) model on a large-scale dataset collected from real-world environments, learning a robust policy that maximizes cover utilization, minimizes threat exposure, and maintains efficient navigation. We demonstrate our method's capabilities on a physical Jackal robot, showing extensive experiments across diverse terrains. These experiments demonstrate EnCoMP's superior performance compared to state-of-the-art methods, achieving a 95% success rate, 85% cover utilization, and reducing threat exposure to 10.5%, while significantly outperforming baselines in navigation efficiency and robustness.
翻译:在复杂环境中运行的自主机器人面临着一个关键挑战:识别并利用环境掩体进行隐蔽导航,以最小化对潜在威胁的暴露。我们提出了EnCoMP,一种增强型导航框架,它集成了离线强化学习与我们新颖的自适应威胁感知可见性估计(ATAVE)算法,使机器人能够在多样化的户外环境中进行隐蔽且高效的导航。ATAVE是我们设计的一种动态概率威胁建模技术,用于持续实时评估和缓解潜在威胁,通过适应不断变化的环境和威胁条件来增强机器人的隐蔽导航能力。此外,我们的方法从LiDAR点云生成高保真度的多地图表示,包括掩体地图、潜在威胁地图、高度地图和目标地图,从而提供对环境的全面理解。这些多地图提供了详细的环境洞察,有助于制定战略性导航决策。目标地图编码了到目标位置的相对距离和方向,引导机器人的导航。我们在从真实世界环境中收集的大规模数据集上训练了一个保守Q学习(CQL)模型,学习到一个鲁棒的策略,该策略最大化掩体利用率、最小化威胁暴露并保持高效的导航。我们在物理Jackal机器人上展示了我们方法的能力,并进行了跨多种地形的广泛实验。这些实验证明了EnCoMP相较于最先进方法的优越性能,实现了95%的成功率、85%的掩体利用率,并将威胁暴露降低至10.5%,同时在导航效率和鲁棒性方面显著优于基线方法。