The growing focus on indoor robot navigation utilizing wireless signals has stemmed from the capability of these signals to capture high-resolution angular and temporal measurements. However, employing end-to-end generic reinforcement learning (RL) for wireless indoor navigation (WIN) in initially unknown environments remains a significant challenge, due to its limited generalization ability and poor sample efficiency. At the same time, purely model-based solutions, based on radio frequency propagation, are simple and generalizable, but unable to find optimal decisions in complex environments. This work proposes a novel physics-informed RL (PIRL) were a standard distance-to-target-based cost along with physics-informed terms on the optimal trajectory. The proposed PIRL is evaluated using a wireless digital twin (WDT) built upon simulations of a large class of indoor environments from the AI Habitat dataset augmented with electromagnetic radiation (EM) simulation for wireless signals. It is shown that the PIRL significantly outperforms both standard RL and purely physics-based solutions in terms of generalizability and performance. Furthermore, the resulting PIRL policy is explainable in that it is empirically consistent with the physics heuristic.
翻译:利用无线信号进行室内机器人导航的研究日益受到关注,因为这类信号能够捕获高分辨率的角域与时域测量信息。然而,在初始未知环境中采用端到端通用强化学习方法进行无线室内导航仍面临重大挑战,主要受限于其有限的泛化能力和较差的样本效率。与此同时,基于射频传播的纯模型驱动方案虽具备简单性和泛化优势,却难以在复杂环境中找到最优决策。本文提出一种新型物理信息强化学习方法,在标准基于目标距离的代价函数基础上,引入与最优轨迹相关的物理信息项。该方法利用无线数字孪生进行评估,该数字孪生基于AI Habitat数据集中大规模室内环境仿真构建,并融合了无线信号的电磁辐射仿真。研究表明,所提出的物理信息强化学习在泛化能力和性能上均显著优于标准强化学习及纯物理驱动方案。此外,该方法生成的策略具有可解释性,其决策结果与物理启发式规则在经验层面保持一致。