In today's world, the rapid expansion of IoT networks and the proliferation of smart devices in our daily lives, have resulted in the generation of substantial amounts of heterogeneous data. These data forms a stream which requires special handling. To handle this data effectively, advanced data processing technologies are necessary to guarantee the preservation of both privacy and efficiency. Federated learning emerged as a distributed learning method that trains models locally and aggregates them on a server to preserve data privacy. This paper showcases two illustrative scenarios that highlight the potential of federated learning (FL) as a key to delivering efficient and privacy-preserving machine learning within IoT networks. We first give the mathematical foundations for key aggregation algorithms in federated learning, i.e., FedAvg and FedProx. Then, we conduct simulations, using Flower Framework, to show the \textit{efficiency} of these algorithms by training deep neural networks on common datasets and show a comparison between the accuracy and loss metrics of FedAvg and FedProx. Then, we present the results highlighting the trade-off between maintaining privacy versus accuracy via simulations - involving the implementation of the differential privacy (DP) method - in Pytorch and Opacus ML frameworks on common FL datasets and data distributions for both FedAvg and FedProx strategies.
翻译:当今世界,物联网网络的快速扩展及智能设备在日常生活中的普及,催生了大量异构数据的生成。这些数据构成需要特殊处理的数据流。为有效处理此类数据,必须采用先进的数据处理技术,以确保隐私和效率的双重保障。联邦学习作为一种分布式学习方法应运而生,该方法在本地训练模型并在服务器端聚合,以保护数据隐私。本文通过两个示例场景,展示联邦学习(FL)作为物联网网络中实现高效且保护隐私的机器学习方案的关键潜力。我们首先给出联邦学习中核心聚合算法(即FedAvg和FedProx)的数学基础。随后,利用Flower框架进行仿真,通过训练深度神经网络在通用数据集上展示这些算法的效率,并对比FedAvg与FedProx的精度与损失指标。接着,我们展示仿真结果——基于Pytorch和Opacus机器学习框架,在通用FL数据集与数据分布上针对FedAvg和FedProx两种策略实施差分隐私(DP)方法——重点揭示隐私保护与精度之间的权衡关系。