In the last few years, the research interest in Vision-and-Language Navigation (VLN) has grown significantly. VLN is a challenging task that involves an agent following human instructions and navigating in a previously unknown environment to reach a specified goal. Recent work in literature focuses on different ways to augment the available datasets of instructions for improving navigation performance by exploiting synthetic training data. In this work, we propose AIGeN, a novel architecture inspired by Generative Adversarial Networks (GANs) that produces meaningful and well-formed synthetic instructions to improve navigation agents' performance. The model is composed of a Transformer decoder (GPT-2) and a Transformer encoder (BERT). During the training phase, the decoder generates sentences for a sequence of images describing the agent's path to a particular point while the encoder discriminates between real and fake instructions. Experimentally, we evaluate the quality of the generated instructions and perform extensive ablation studies. Additionally, we generate synthetic instructions for 217K trajectories using AIGeN on Habitat-Matterport 3D Dataset (HM3D) and show an improvement in the performance of an off-the-shelf VLN method. The validation analysis of our proposal is conducted on REVERIE and R2R and highlights the promising aspects of our proposal, achieving state-of-the-art performance.
翻译:近年来,视觉语言导航(VLN)的研究兴趣显著增长。VLN是一项具有挑战性的任务,涉及智能体遵循人类指令,在未知环境中导航以到达指定目标。近期文献工作侧重于通过利用合成训练数据,以不同方式扩充现有指令数据集来提升导航性能。本文提出了一种受生成对抗网络(GANs)启发的新颖架构AIGeN,该架构能够生成有意义且结构良好的合成指令,从而提高导航智能体的性能。该模型由Transformer解码器(GPT-2)和Transformer编码器(BERT)组成。在训练阶段,解码器针对描述智能体到达特定点路径的图像序列生成句子,而编码器则区分真实指令与虚假指令。实验方面,我们评估了生成指令的质量,并进行了广泛的消融研究。此外,我们使用AIGeN在Habitat-Matterport 3D数据集(HM3D)上为217K条轨迹生成了合成指令,并展示了其对现成VLN方法性能的提升。我们的验证分析在REVERIE和R2R数据集上进行,突出了所提方案的有前景之处,并达到了最先进的性能。