The generalization of the end-to-end deep reinforcement learning (DRL) for object-goal visual navigation is a long-standing challenge since object classes and placements vary in new test environments. Learning domain-independent visual representation is critical for enabling the trained DRL agent with the ability to generalize to unseen scenes and objects. In this letter, a target-directed attention network (TDANet) is proposed to learn the end-to-end object-goal visual navigation policy with zero-shot ability. TDANet features a novel target attention (TA) module that learns both the spatial and semantic relationships among objects to help TDANet focus on the most relevant observed objects to the target. With the Siamese architecture (SA) design, TDANet distinguishes the difference between the current and target states and generates the domain-independent visual representation. To evaluate the navigation performance of TDANet, extensive experiments are conducted in the AI2-THOR embodied AI environment. The simulation results demonstrate a strong generalization ability of TDANet to unseen scenes and target objects, with higher navigation success rate (SR) and success weighted by length (SPL) than other state-of-the-art models.
翻译:端到端深度强化学习在物体目标视觉导航任务中的泛化能力始终面临挑战,因为新测试环境中的物体类别与空间布局存在显著差异。学习与领域无关的视觉表征对于赋予训练后的深度强化学习智能体泛化至未见场景与物体的能力至关重要。本文提出了一种目标导向注意力网络(TDANet),用于学习具有零样本能力的端到端物体目标视觉导航策略。TDANet创新性地设计了目标注意力模块,该模块可同时学习物体间的空间与语义关联,帮助网络聚焦与目标最相关的观测物体。通过孪生网络架构的设计,TDANet能够区分当前状态与目标状态的差异,并生成领域无关的视觉表征。为评估TDANet的导航性能,我们在AI2-THOR具身人工智能环境中开展了大量实验。仿真结果表明,相较于其他先进模型,TDANet对未见场景与目标物体展现出强大的泛化能力,并取得了更高的导航成功率与路径长度加权成功率。