This paper considers the secure aggregation problem for federated learning under an information theoretic cryptographic formulation, where distributed training nodes (referred to as users) train models based on their own local data and a curious-but-honest server aggregates the trained models without retrieving other information about users' local data. Secure aggregation generally contains two phases, namely key sharing phase and model aggregation phase. Due to the common effect of user dropouts in federated learning, the model aggregation phase should contain two rounds, where in the first round the users transmit masked models and, in the second round, according to the identity of surviving users after the first round, these surviving users transmit some further messages to help the server decrypt the sum of users' trained models. The objective of the considered information theoretic formulation is to characterize the capacity region of the communication rates in the two rounds from the users to the server in the model aggregation phase, assuming that key sharing has already been performed offline in prior. In this context, Zhao and Sun completely characterized the capacity region under the assumption that the keys can be arbitrary random variables. More recently, an additional constraint, known as "uncoded groupwise keys," has been introduced. This constraint entails the presence of multiple independent keys within the system, with each key being shared by precisely S users. The capacity region for the information-theoretic secure aggregation problem with uncoded groupwise keys was established in our recent work subject to the condition S > K - U, where K is the number of total users and U is the designed minimum number of surviving users. In this paper we fully characterize of the the capacity region for this problem by proposing a new converse bound and an achievable scheme.
翻译:本文从信息论密码学角度研究了联邦学习中的安全聚合问题,其中分布式训练节点(称为用户)基于自身本地数据训练模型,而好奇心但诚实的服务器在聚合训练模型时不获取用户本地数据的其他信息。安全聚合通常包含两个阶段:密钥共享阶段和模型聚合阶段。由于联邦学习中常见的用户退出效应,模型聚合阶段应包含两轮:第一轮中用户传输掩码模型,第二轮中根据第一轮后幸存用户的身份,这些幸存用户传输额外消息以帮助服务器解密用户训练模型的总和。所考虑的信息论公式的目标是在假设密钥共享已事先离线完成的情况下,刻画模型聚合阶段两轮中用户到服务器通信速率的容量域。在此背景下,Zhao和Sun在假设密钥可为任意随机变量时完整刻画了该容量域。最近,引入了一项额外约束条件,称为"无编码分组密钥"。该约束要求系统中存在多个独立密钥,每个密钥恰好由S个用户共享。我们在近期工作中建立了满足条件S > K - U(其中K为用户总数,U为设计的最少幸存用户数)时无编码分组密钥信息论安全聚合问题的容量域。本文通过提出新的逆向界和可实现方案,完整刻画了该问题的容量域。