Split learning (SL) aims to protect user data privacy by distributing deep models between client-server and keeping private data locally. In SL training with multiple clients, the local model weights are shared among the clients for local model update. This paper first reveals data privacy leakage exacerbated from local weight sharing among the clients in SL through model inversion attacks. Then, to reduce the data privacy leakage issue, we propose and analyze privacy-enhanced SL (P-SL) (or SL without local weight sharing). We further propose parallelized P-SL to expedite the training process by duplicating multiple server-side model instances without compromising accuracy. Finally, we explore P-SL with late participating clients and devise a server-side cache-based training method to address the forgetting phenomenon in SL when late clients join. Experimental results demonstrate that P-SL helps reduce up to 50% of client-side data leakage, which essentially achieves a better privacy-accuracy trade-off than the current trend by using differential privacy mechanisms. Moreover, P-SL and its cache-based version achieve comparable accuracy to baseline SL under various data distributions, while cost less computation and communication. Additionally, caching-based training in P-SL mitigates the negative effect of forgetting, stabilizes the learning, and enables practical and low-complexity training in a dynamic environment with late-arriving clients.
翻译:分割学习(SL)旨在通过将深度模型分布在客户端-服务器之间,并将私有数据保留在本地,从而保护用户数据隐私。在涉及多个客户端的SL训练中,本地模型权重在客户端之间共享以进行本地模型更新。本文首先通过模型逆向攻击揭示了SL中客户端间本地权重共享加剧了数据隐私泄露。然后,为减轻数据隐私泄露问题,我们提出并分析了隐私增强型分割学习(P-SL),即无本地权重共享的分割学习。我们进一步提出并行化P-SL,通过复制多个服务器端模型实例来加速训练过程,同时不牺牲准确性。最后,我们探索了面向迟到参与客户端的P-SL,并设计了一种基于服务器端缓存的训练方法,以解决迟到客户端加入时SL中的遗忘现象。实验结果表明,P-SL有助于减少高达50%的客户端数据泄露,这本质上比当前使用差分隐私机制的方法实现了更好的隐私-准确性权衡。此外,P-SL及其基于缓存的版本在各种数据分布下均能达到与基准SL相当的准确性,同时计算和通信成本更低。另外,P-SL中的基于缓存训练减轻了遗忘的负面影响,稳定了学习过程,并在存在迟到客户端的动态环境中实现了低复杂度的实用化训练。