Artificial intelligence (AI) models are increasingly used in the medical domain. However, as medical data is highly sensitive, special precautions to ensure its protection are required. The gold standard for privacy preservation is the introduction of differential privacy (DP) to model training. Prior work indicates that DP has negative implications on model accuracy and fairness, which are unacceptable in medicine and represent a main barrier to the widespread use of privacy-preserving techniques. In this work, we evaluated the effect of privacy-preserving training of AI models for chest radiograph diagnosis regarding accuracy and fairness compared to non-private training. For this, we used a large dataset (N=193,311) of high quality clinical chest radiographs, which were retrospectively collected and manually labeled by experienced radiologists. We then compared non-private deep convolutional neural networks (CNNs) and privacy-preserving (DP) models with respect to privacy-utility trade-offs measured as area under the receiver-operator-characteristic curve (AUROC), and privacy-fairness trade-offs, measured as Pearson's r or Statistical Parity Difference. We found that the non-private CNNs achieved an average AUROC score of 0.90 +- 0.04 over all labels, whereas the DP CNNs with a privacy budget of epsilon=7.89 resulted in an AUROC of 0.87 +- 0.04, i.e., a mere 2.6% performance decrease compared to non-private training. Furthermore, we found the privacy-preserving training not to amplify discrimination against age, sex or co-morbidity. Our study shows that -- under the challenging realistic circumstances of a real-life clinical dataset -- the privacy-preserving training of diagnostic deep learning models is possible with excellent diagnostic accuracy and fairness.
翻译:人工智能(AI)模型在医疗领域中的应用日益广泛。然而,由于医疗数据高度敏感,需要采取特殊保护措施以确保其安全性。隐私保护的金标准是在模型训练中引入差分隐私。先前研究表明,差分隐私会对模型准确性和公平性产生负面影响,这在医学领域是不可接受的,并成为隐私保护技术广泛使用的主要障碍。本研究评估了在胸部X光片诊断中,与无隐私训练相比,隐私保护AI模型训练对准确性和公平性的影响。为此,我们使用了大规模(N=193,311)高质量临床胸部X光片数据集,这些数据经过回顾性收集并由经验丰富的放射科医生手动标注。随后,我们比较了无隐私深度卷积神经网络与隐私保护差分隐私模型在隐私-效用权衡(以受试者工作特征曲线下面积AUROC衡量)和隐私-公平性权衡(以Pearson相关系数或统计奇偶性差异衡量)方面的表现。研究发现,无隐私CNN在所有标签上的平均AUROC得分为0.90 ± 0.04,而隐私预算ε=7.89的DP-CNN的AUROC得分为0.87 ± 0.04,即与无隐私训练相比性能仅下降2.6%。此外,我们发现隐私保护训练并未加剧对年龄、性别或合并症的歧视。本研究证明——在真实临床数据集的具有挑战性的现实条件下——对诊断深度学习模型进行隐私保护训练是可行的,且能够保持优异的诊断准确性和公平性。