Split learning (SL) aims to protect user data privacy by distributing deep models between client-server and keeping private data locally. In SL training with multiple clients, the local model weights are shared among the clients for local model update. This paper first reveals data privacy leakage exacerbated from local weight sharing among the clients in SL through model inversion attacks. Then, to reduce the data privacy leakage issue, we propose and analyze privacy-enhanced SL (P-SL) (or SL without local weight sharing). We further propose parallelized P-SL to expedite the training process by duplicating multiple server-side model instances without compromising accuracy. Finally, we explore P-SL with late participating clients and devise a server-side cache-based training method to address the forgetting phenomenon in SL when late clients join. Experimental results demonstrate that P-SL helps reduce up to 50% of client-side data leakage, which essentially achieves a better privacy-accuracy trade-off than the current trend by using differential privacy mechanisms. Moreover, P-SL and its cache-based version achieve comparable accuracy to baseline SL under various data distributions, while cost less computation and communication. Additionally, caching-based training in P-SL mitigates the negative effect of forgetting, stabilizes the learning, and enables practical and low-complexity training in a dynamic environment with late-arriving clients.
翻译:拆分学习(SL)旨在通过将深度模型分布于客户端与服务器之间并保持私有数据本地化来保护用户数据隐私。在多客户端SL训练中,本地模型权重会在客户端间共享以更新本地模型。本文首先揭示了SL中客户端间本地权重共享会通过模型反演攻击加剧数据隐私泄露。为缓解数据隐私泄露问题,我们提出并分析了隐私增强型SL(P-SL)(即无需本地权重共享的SL)。我们进一步提出并行化P-SL,通过复制多个服务器端模型实例来加速训练过程且不损失精度。最后,我们探索了P-SL在延迟参与客户端场景下的应用,并设计了一种基于服务器缓存的训练方法,以解决延迟客户端加入时SL中的遗忘现象。实验结果表明,P-SL能帮助减少高达50%的客户端数据泄露,相比当前使用差分隐私机制的趋势,本质上实现了更优的隐私-精度权衡。此外,P-SL及其缓存版本在不同数据分布下达到了与基线SL相当的精度,同时降低了计算与通信开销。基于缓存的P-SL训练还能减轻遗忘的负面影响,稳定学习过程,并在延迟客户端到达的动态环境中实现实用且低复杂度的训练。