Federated Learning (FL) enables distributed participants (e.g., mobile devices) to train a global model without sharing data directly to a central server. Recent studies have revealed that FL is vulnerable to gradient inversion attack (GIA), which aims to reconstruct the original training samples and poses high risk against the privacy of clients in FL. However, most existing GIAs necessitate control over the server and rely on strong prior knowledge including batch normalization and data distribution information. In this work, we propose Client-side poisoning Gradient Inversion (CGI), which is a novel attack method that can be launched from clients. For the first time, we show the feasibility of a client-side adversary with limited knowledge being able to recover the training samples from the aggregated global model. We take a distinct approach in which the adversary utilizes a malicious model that amplifies the loss of a specific targeted class of interest. When honest clients employ the poisoned global model, the gradients of samples belonging to the targeted class are magnified, making them the dominant factor in the aggregated update. This enables the adversary to effectively reconstruct the private input belonging to other clients using the aggregated update. In addition, our CGI also features its ability to remain stealthy against Byzantine-robust aggregation rules (AGRs). By optimizing malicious updates and blending benign updates with a malicious replacement vector, our method remains undetected by these defense mechanisms. To evaluate the performance of CGI, we conduct experiments on various benchmark datasets, considering representative Byzantine-robust AGRs, and exploring diverse FL settings with different levels of adversary knowledge about the data. Our results demonstrate that CGI consistently and successfully extracts training input in all tested scenarios.
翻译:联邦学习(FL)使分布式参与者(如移动设备)能够在不直接向中央服务器共享数据的情况下训练全局模型。近期研究表明,联邦学习易受梯度反演攻击(GIA),该攻击旨在重建原始训练样本,对联邦学习客户端的隐私构成高风险。然而,现有大多数GIA需控制服务器,且依赖批归一化和数据分布信息等强先验知识。在本工作中,我们提出客户端侧投毒梯度反演(CGI),这是一种可从客户端发起的全新攻击方法。我们首次展示了具备有限知识的客户端侧攻击者能够从聚合全局模型中恢复训练样本的可行性。我们采用独特方法:攻击者利用恶意模型放大特定目标类别的损失。当诚实客户端使用中毒全局模型时,属于目标类别的样本梯度被放大,使其成为聚合更新中的主导因素。这使得攻击者能够有效利用聚合更新重建其他客户端所属的私有输入。此外,CGI还具有抵御拜占庭鲁棒聚合规则(AGRs)的隐蔽能力。通过优化恶意更新并用恶意替换向量混合良性更新,我们的方法可绕过这些防御机制的检测。为评估CGI性能,我们在多种基准数据集上进行实验,考虑代表性拜占庭鲁棒AGRs,并探索攻击者具备不同数据知识水平的多样化FL设置。结果表明,CGI在所有测试场景中均能持续成功提取训练输入。