This paper explores how to customize time series classification (TSC) methods with the help of external data in a privacy-preserving federated learning (FL) scenario. To the best of our knowledge, we are the first to study on this essential topic. Achieving this goal requires us to seamlessly integrate the techniques from multiple fields including Data Mining, Machine Learning, and Security. In this paper, we systematically investigate existing TSC solutions for the centralized scenario and propose FedST, a novel FL-enabled TSC framework based on a shapelet transformation method. We recognize the federated shapelet search step as the kernel of FedST. Thus, we design a basic protocol for the FedST kernel that we prove to be secure and accurate. However, we identify that the basic protocol suffers from efficiency bottlenecks and the centralized acceleration techniques lose their efficacy due to the security issues. To speed up the federated protocol with security guarantee, we propose several optimizations tailored for the FL setting. Our theoretical analysis shows that the proposed methods are secure and more efficient. We conduct extensive experiments using both synthetic and real-world datasets. Empirical results show that our FedST solution is effective in terms of TSC accuracy, and the proposed optimizations can achieve three orders of magnitude of speedup.
翻译:本文探索如何在隐私保护的联邦学习场景中借助外部数据定制时间序列分类方法。据我们所知,这是首次针对这一关键问题展开研究。实现该目标需要无缝融合数据挖掘、机器学习与安全等多个领域的技术。本文系统调研了现有集中式场景下的时间序列分类解决方案,并提出FedST——一种基于形状变换方法的新型联邦学习时间序列分类框架。我们将联邦形状搜索步骤视为FedST的核心,因此为其设计了基础协议,并证明该协议兼具安全性与准确性。然而,我们发现基础协议存在效率瓶颈,且由于安全问题,集中式加速技术无法有效应用。为在保障安全的前提下加速联邦协议,我们提出了多项针对联邦学习场景的优化方案。理论分析表明,所提方法既安全又高效。通过使用合成数据集和真实数据集开展大量实验,实证结果显示:FedST在时间序列分类精度上表现优异,且所提优化方案可实现三个数量级的加速效果。