Most existing works solving Room-to-Room VLN problem only utilize RGB images and do not consider local context around candidate views, which lack sufficient visual cues about surrounding environment. Moreover, natural language contains complex semantic information thus its correlations with visual inputs are hard to model merely with cross attention. In this paper, we propose GeoVLN, which learns Geometry-enhanced visual representation based on slot attention for robust Visual-and-Language Navigation. The RGB images are compensated with the corresponding depth maps and normal maps predicted by Omnidata as visual inputs. Technically, we introduce a two-stage module that combine local slot attention and CLIP model to produce geometry-enhanced representation from such input. We employ V&L BERT to learn a cross-modal representation that incorporate both language and vision informations. Additionally, a novel multiway attention module is designed, encouraging different phrases of input instruction to exploit the most related features from visual input. Extensive experiments demonstrate the effectiveness of our newly designed modules and show the compelling performance of the proposed method.
翻译:现有解决房间到房间视觉-语言导航问题的大多数工作仅利用RGB图像,未考虑候选视图周围的局部上下文,导致缺乏关于周围环境的足够视觉线索。此外,自然语言包含复杂的语义信息,其与视觉输入的相关性难以仅通过交叉注意力进行建模。本文提出GeoVLN方法,基于槽注意力学习几何增强视觉表示,以实现鲁棒的视觉-语言导航。该方法通过Omnidata预测的对应深度图和法向图对RGB图像进行补偿,作为视觉输入。在技术上,我们引入一个两阶段模块,将局部槽注意力与CLIP模型结合,从此类输入中生成几何增强表示。采用V&L BERT学习融合语言与视觉信息的跨模态表示。此外,设计了一种新颖的多路注意力模块,促使输入指令的不同短语从视觉输入中挖掘最相关的特征。大量实验证明了所设计模块的有效性,并展示了所提方法的卓越性能。