Federated learning enables multiple decentralized clients to learn collaboratively without sharing the local training data. However, the expensive annotation cost to acquire data labels on local clients remains an obstacle in utilizing local data. In this paper, we propose a federated active learning paradigm to efficiently learn a global model with limited annotation budget while protecting data privacy in a decentralized learning way. The main challenge faced by federated active learning is the mismatch between the active sampling goal of the global model on the server and that of the asynchronous local clients. This becomes even more significant when data is distributed non-IID across local clients. To address the aforementioned challenge, we propose Knowledge-Aware Federated Active Learning (KAFAL), which consists of Knowledge-Specialized Active Sampling (KSAS) and Knowledge-Compensatory Federated Update (KCFU). KSAS is a novel active sampling method tailored for the federated active learning problem. It deals with the mismatch challenge by sampling actively based on the discrepancies between local and global models. KSAS intensifies specialized knowledge in local clients, ensuring the sampled data to be informative for both the local clients and the global model. KCFU, in the meantime, deals with the client heterogeneity caused by limited data and non-IID data distributions. It compensates for each client's ability in weak classes by the assistance of the global model. Extensive experiments and analyses are conducted to show the superiority of KSAS over the state-of-the-art active learning methods and the efficiency of KCFU under the federated active learning framework.
翻译:联邦学习使多个分散的客户端无需共享局部训练数据即可协作学习。然而,在本地客户端上获取数据标签所需的高昂标注成本仍然是利用局部数据的障碍。本文提出了一种联邦主动学习范式,以在分散学习方式下高效学习全局模型,同时保护数据隐私,并限制标注预算。联邦主动学习面临的主要挑战是服务器上全局模型的主动采样目标与异步局部客户端的采样目标之间的不匹配。当数据以非独立同分布方式分布在各个本地客户端时,这一问题更为显著。为应对上述挑战,我们提出知识感知联邦主动学习(KAFAL),该方法包含知识专化主动采样(KSAS)和知识补偿联邦更新(KCFU)两个模块。KSAS是针对联邦主动学习问题设计的新型主动采样方法,通过基于局部与全局模型差异进行主动采样来应对不匹配挑战。KSAS强化局部客户端中的专化知识,确保采样数据对局部客户端和全局模型均具有信息量。同时,KCFU处理由有限数据和非独立同分布数据分布导致的客户端异质性,借助全局模型补偿每个客户端在弱类上的能力。通过大量实验与分析,验证了KSAS相较于当前最先进主动学习方法的优越性,以及KCFU在联邦主动学习框架下的有效性。