The prominence of embodied Artificial Intelligence (AI), which empowers robots to navigate, perceive, and engage within virtual environments, has attracted significant attention, owing to the remarkable advancements in computer vision and large language models. Privacy emerges as a pivotal concern within the realm of embodied AI, as the robot accesses substantial personal information. However, the issue of privacy leakage in embodied AI tasks, particularly in relation to reinforcement learning algorithms, has not received adequate consideration in research. This paper aims to address this gap by proposing an attack on the value-based algorithm and the gradient-based algorithm, utilizing gradient inversion to reconstruct states, actions, and supervision signals. The choice of using gradients for the attack is motivated by the fact that commonly employed federated learning techniques solely utilize gradients computed based on private user data to optimize models, without storing or transmitting the data to public servers. Nevertheless, these gradients contain sufficient information to potentially expose private data. To validate our approach, we conduct experiments on the AI2THOR simulator and evaluate our algorithm on active perception, a prevalent task in embodied AI. The experimental results demonstrate the effectiveness of our method in successfully reconstructing all information from the data across 120 room layouts.
翻译:具身人工智能(AI)的兴起——使机器人能够在虚拟环境中导航、感知和交互——因其在计算机视觉和大语言模型领域的显著进展而备受关注。隐私成为具身AI领域的一个关键问题,因为机器人会接触到大量个人信息。然而,在具身AI任务中,尤其是与强化学习算法相关的隐私泄露问题,尚未在研究中得到充分重视。本文旨在填补这一空白,提出一种针对基于价值的算法和基于梯度的算法的攻击方法,利用梯度反演来重建状态、动作和监督信号。选择使用梯度进行攻击的动机在于,常用的联邦学习技术仅利用基于私有用户数据计算的梯度来优化模型,而不存储或传输数据到公共服务器。然而,这些梯度包含足够的信息,可能暴露私有数据。为了验证我们的方法,我们在AI2THOR模拟器上进行实验,并针对具身AI中常见任务——主动感知——评估我们的算法。实验结果证明了我们的方法在120种房间布局中成功从数据中重建所有信息的有效性。