We introduce a novel speaker model \textsc{Kefa} for navigation instruction generation. The existing speaker models in Vision-and-Language Navigation suffer from the large domain gap of vision features between different environments and insufficient temporal grounding capability. To address the challenges, we propose a Knowledge Refinement Module to enhance the feature representation with external knowledge facts, and an Adaptive Temporal Alignment method to enforce fine-grained alignment between the generated instructions and the observation sequences. Moreover, we propose a new metric SPICE-D for navigation instruction evaluation, which is aware of the correctness of direction phrases. The experimental results on R2R and UrbanWalk datasets show that the proposed KEFA speaker achieves state-of-the-art instruction generation performance for both indoor and outdoor scenes.
翻译:我们提出了一种新颖的说话模型Kefa,用于导航指令生成。现有的视觉-语言导航中的说话模型面临不同环境间视觉特征域差异大以及时间对齐能力不足的问题。为应对这些挑战,我们提出了一种知识精炼模块,利用外部知识事实增强特征表示,并设计了一种自适应时间对齐方法,以实现生成的指令与观测序列之间的细粒度对齐。此外,我们还提出了一种新的导航指令评估指标SPICE-D,能够感知方向短语的准确性。在R2R和UrbanWalk数据集上的实验结果表明,所提出的Kefa说话模型在室内和室外场景中均实现了最先进的指令生成性能。