Developing robust artificial intelligence (AI) models that generalize well to unseen datasets is challenging and usually requires large and variable datasets, preferably from multiple institutions. In federated learning (FL), a model is trained collaboratively at numerous sites that hold local datasets without exchanging them. So far, the impact of training strategy, i.e., local versus collaborative, on the diagnostic on-domain and off-domain performance of AI models interpreting chest radiographs has not been assessed. Consequently, using 610,000 chest radiographs from five institutions across the globe, we assessed diagnostic performance as a function of training strategy (i.e., local vs. collaborative), network architecture (i.e., convolutional vs. transformer-based), generalization performance (i.e., on-domain vs. off-domain), imaging finding (i.e., cardiomegaly, pleural effusion, pneumonia, atelectasis, consolidation, pneumothorax, and no abnormality), dataset size (i.e., from n=18,000 to 213,921 radiographs), and dataset diversity. Large datasets not only showed minimal performance gains with FL but, in some instances, even exhibited decreases. In contrast, smaller datasets revealed marked improvements. Thus, on-domain performance was mainly driven by training data size. However, off-domain performance leaned more on training diversity. When trained collaboratively across diverse external institutions, AI models consistently surpassed models trained locally for off-domain tasks, emphasizing FL's potential in leveraging data diversity. In conclusion, FL can bolster diagnostic privacy, reproducibility, and off-domain reliability of AI models and, potentially, optimize healthcare outcomes.
翻译:开发能够良好泛化至未见数据集的鲁棒人工智能(AI)模型具有挑战性,通常需要大规模且多样化的数据集,最好来自多家机构。在联邦学习(FL)中,模型在多个持有本地数据集但无需交换数据的站点协作训练。迄今为止,训练策略(即本地训练与协作训练)对解读胸部X光片的AI模型在域内和域外诊断性能的影响尚未得到评估。因此,利用来自全球五家机构的61万张胸部X光片,我们评估了诊断性能与训练策略(本地vs.协作)、网络架构(卷积vs.基于Transformer)、泛化性能(域内vs.域外)、影像学发现(心脏肥大、胸腔积液、肺炎、肺不张、实变、气胸和无异常)、数据集规模(从18,000张到213,921张X光片)以及数据集多样性的函数关系。大数据集不仅在FL下性能提升微乎其微,某些情况下甚至出现下降;相反,小数据集则展现出显著改善。因此,域内性能主要受训练数据规模驱动;而域外性能更依赖于训练多样性。当跨不同外部机构协作训练时,AI模型在域外任务中始终优于本地训练模型,凸显了FL在利用数据多样性方面的潜力。总之,联邦学习可增强AI模型的诊断隐私性、可重复性及域外可靠性,并有可能优化医疗健康结果。