Federated learning facilitates the collaborative learning of a global model across multiple distributed medical institutions without centralizing data. Nevertheless, the expensive cost of annotation on local clients remains an obstacle to effectively utilizing local data. To mitigate this issue, federated active learning methods suggest leveraging local and global model predictions to select a relatively small amount of informative local data for annotation. However, existing methods mainly focus on all local data sampled from the same domain, making them unreliable in realistic medical scenarios with domain shifts among different clients. In this paper, we make the first attempt to assess the informativeness of local data derived from diverse domains and propose a novel methodology termed Federated Evidential Active Learning (FEAL) to calibrate the data evaluation under domain shift. Specifically, we introduce a Dirichlet prior distribution in both local and global models to treat the prediction as a distribution over the probability simplex and capture both aleatoric and epistemic uncertainties by using the Dirichlet-based evidential model. Then we employ the epistemic uncertainty to calibrate the aleatoric uncertainty. Afterward, we design a diversity relaxation strategy to reduce data redundancy and maintain data diversity. Extensive experiments and analyses are conducted to show the superiority of FEAL over the state-of-the-art active learning methods and the efficiency of FEAL under the federated active learning framework.
翻译:联邦学习允许多个分布式医疗机构在不集中数据的情况下协同训练全局模型。然而,本地客户端高昂的标注成本仍是充分利用本地数据的障碍。为解决这一问题,联邦主动学习方法通过利用本地和全局模型的预测,选择少量信息量丰富的本地数据进行标注。然而,现有方法主要关注同一领域采样的本地数据,在存在不同客户端间领域偏移的现实医疗场景中并不可靠。本文首次尝试评估来自不同领域的本地数据的信息量,并提出一种名为联邦证据主动学习(FEAL)的新方法,以校准领域偏移下的数据评估。具体而言,我们在本地和全局模型中引入狄利克雷先验分布,将预测视为概率单纯形上的分布,并利用基于狄利克雷的证据模型捕获偶然不确定性和认知不确定性。随后,我们利用认知不确定性校准偶然不确定性。此外,我们设计了一种多样性松弛策略以减少数据冗余并保持数据多样性。通过大量实验和分析,证明了FEAL相较于最先进的主动学习方法的优越性,以及其在联邦主动学习框架下的高效性。