Embodied agents in vision navigation coupled with deep neural networks have attracted increasing attention. However, deep neural networks have been shown vulnerable to malicious adversarial noises, which may potentially cause catastrophic failures in Embodied Vision Navigation. Among different adversarial noises, universal adversarial perturbations (UAP), i.e., a constant image-agnostic perturbation applied on every input frame of the agent, play a critical role in Embodied Vision Navigation since they are computation-efficient and application-practical during the attack. However, existing UAP methods ignore the system dynamics of Embodied Vision Navigation and might be sub-optimal. In order to extend UAP to the sequential decision setting, we formulate the disturbed environment under the universal noise $\delta$, as a $\delta$-disturbed Markov Decision Process ($\delta$-MDP). Based on the formulation, we analyze the properties of $\delta$-MDP and propose two novel Consistent Attack methods, named Reward UAP and Trajectory UAP, for attacking Embodied agents, which consider the dynamic of the MDP and calculate universal noises by estimating the disturbed distribution and the disturbed Q function. For various victim models, our Consistent Attack can cause a significant drop in their performance in the PointGoal task in Habitat with different datasets and different scenes. Extensive experimental results indicate that there exist serious potential risks for applying Embodied Vision Navigation methods to the real world.
翻译:与深度神经网络结合的具身视觉导航智能体日益受到关注。然而,深度神经网络已被证实对恶意对抗噪声具有脆弱性,这可能导致具身视觉导航出现灾难性故障。在不同对抗噪声中,通用对抗扰动(UAP)(即对智能体每个输入帧施加的恒定且与图像无关的扰动)在具身视觉导航中发挥关键作用,因其在攻击过程中具有计算高效性和应用实用性。然而,现有UAP方法忽略了具身视觉导航的系统动力学特性,可能仅达到次优效果。为将UAP扩展至序贯决策场景,我们首先将通用噪声$\delta$作用下的扰动环境形式化为$\delta$-扰动马尔可夫决策过程($\delta$-MDP)。基于该形式化,我们分析了$\delta$-MDP的属性,并提出了两种新颖的一致性攻击方法——奖赏UAP和轨迹UAP,用于攻击具身智能体。这些方法充分考虑MDP的动力学特性,通过估计扰动分布和扰动Q函数来计算通用噪声。针对各类受害者模型,我们的一致性攻击方法在Habitat平台中不同数据集和场景下的PointGoal任务中,均能导致其性能显著下降。大量实验结果表明,将具身视觉导航方法应用于真实世界存在严重的潜在风险。