Motivated by the advancing computational capacity of distributed end-user equipments (UEs), as well as the increasing concerns about sharing private data, there has been considerable recent interest in machine learning (ML) and artificial intelligence (AI) that can be processed on on distributed UEs. Specifically, in this paradigm, parts of an ML process are outsourced to multiple distributed UEs, and then the processed ML information is aggregated on a certain level at a central server, which turns a centralized ML process into a distributed one, and brings about significant benefits. However, this new distributed ML paradigm raises new risks of privacy and security issues. In this paper, we provide a survey of the emerging security and privacy risks of distributed ML from a unique perspective of information exchange levels, which are defined according to the key steps of an ML process, i.e.: i) the level of preprocessed data, ii) the level of learning models, iii) the level of extracted knowledge and, iv) the level of intermediate results. We explore and analyze the potential of threats for each information exchange level based on an overview of the current state-of-the-art attack mechanisms, and then discuss the possible defense methods against such threats. Finally, we complete the survey by providing an outlook on the challenges and possible directions for future research in this critical area.
翻译:由于分布式终端设备(UEs)计算能力的不断提升,以及人们对共享私人数据的日益关注,近年来,可在分布式终端设备上处理的机器学习(ML)和人工智能(AI)引起了广泛兴趣。具体而言,在这一范式中,ML过程的某些部分被外包给多个分布式终端设备,随后处理后的ML信息在中央服务器上以特定层级进行聚合,从而将集中式ML过程转变为分布式过程,并带来显著优势。然而,这种新型分布式ML范式也引发了新的隐私与安全风险。本文从信息交换层级的独特视角出发(根据ML过程的关键步骤定义,即:i) 预处理数据层级,ii) 学习模型层级,iii) 提取知识层级,以及iv) 中间结果层级),对分布式ML中新兴的安全与隐私风险进行了综述。我们基于对当前最新攻击机制的概述,探索并分析了每个信息交换层级的潜在威胁,随后讨论了针对此类威胁可能的防御方法。最后,本文对这一关键领域未来的研究挑战与可能方向进行了展望,从而完成综述。