Robust and reliable anonymization of chest radiographs constitutes an essential step before publishing large datasets of such for research purposes. The conventional anonymization process is carried out by obscuring personal information in the images with black boxes and removing or replacing meta-information. However, such simple measures retain biometric information in the chest radiographs, allowing patients to be re-identified by a linkage attack. Therefore, there is an urgent need to obfuscate the biometric information appearing in the images. We propose the first deep learning-based approach (PriCheXy-Net) to targetedly anonymize chest radiographs while maintaining data utility for diagnostic and machine learning purposes. Our model architecture is a composition of three independent neural networks that, when collectively used, allow for learning a deformation field that is able to impede patient re-identification. Quantitative results on the ChestX-ray14 dataset show a reduction of patient re-identification from 81.8% to 57.7% (AUC) after re-training with little impact on the abnormality classification performance. This indicates the ability to preserve underlying abnormality patterns while increasing patient privacy. Lastly, we compare our proposed anonymization approach with two other obfuscation-based methods (Privacy-Net, DP-Pix) and demonstrate the superiority of our method towards resolving the privacy-utility trade-off for chest radiographs.
翻译:胸部X光片的鲁棒且可靠的匿名化是将此类大型数据集用于研究目的之前的关键步骤。传统的匿名化过程通过使用黑框遮蔽图像中的个人信息,并移除或替换元信息来执行。然而,此类简单措施会保留胸片中的生物特征信息,使得患者可能通过链接攻击被重新识别。因此,迫切需要模糊化图像中出现的生物特征信息。我们提出了首个基于深度学习的方法(PriCheXy-Net),旨在在保持诊断和机器学习用途的数据实用性的同时,对胸片进行目标性匿名化。我们的模型架构由三个独立神经网络组合而成,当共同使用时,能够学习一个变形场,从而阻止患者重新识别。在ChestX-ray14数据集上的定量结果显示,重新训练后患者重新识别率从81.8%降至57.7%(AUC),而对异常分类性能影响极小。这表明该方法能够保留潜在的异常模式,同时增强患者隐私。最后,我们将所提出的匿名化方法与另外两种基于模糊化的方法(Privacy-Net、DP-Pix)进行比较,并证明我们的方法在解决胸片隐私-实用性权衡方面的优越性。