Machine learning (ML) can help fight pandemics like COVID-19 by enabling rapid screening of large volumes of images. To perform data analysis while maintaining patient privacy, we create ML models that satisfy Differential Privacy (DP). Previous works exploring private COVID-19 models are in part based on small datasets, provide weaker or unclear privacy guarantees, and do not investigate practical privacy. We suggest improvements to address these open gaps. We account for inherent class imbalances and evaluate the utility-privacy trade-off more extensively and over stricter privacy budgets. Our evaluation is supported by empirically estimating practical privacy through black-box Membership Inference Attacks (MIAs). The introduced DP should help limit leakage threats posed by MIAs, and our practical analysis is the first to test this hypothesis on the COVID-19 classification task. Our results indicate that needed privacy levels might differ based on the task-dependent practical threat from MIAs. The results further suggest that with increasing DP guarantees, empirical privacy leakage only improves marginally, and DP therefore appears to have a limited impact on practical MIA defense. Our findings identify possibilities for better utility-privacy trade-offs, and we believe that empirical attack-specific privacy estimation can play a vital role in tuning for practical privacy.
翻译:机器学习(ML)可通过快速筛查大量图像来助力应对COVID-19等疫情。为在数据分析过程中保护患者隐私,我们构建了满足差分隐私(DP)的机器学习模型。此前关于隐私COVID-19模型的研究部分基于小型数据集,提供的隐私保障较弱或定义模糊,且未深入探讨实际隐私问题。我们提出改进方案以填补这些空白:考虑了固有的类别不平衡问题,在更严格的隐私预算下更全面地评估效用-隐私权衡。我们的评估通过黑盒成员推断攻击(MIAs)对实际隐私进行经验估计来支持。引入的DP应有助于限制MIAs带来的泄漏威胁,而我们是首个在COVID-19分类任务中检验此假设的实践性研究。结果表明所需隐私水平可能取决于MIAs基于任务的实际威胁程度。结果进一步表明,随着DP保障增强,经验性隐私泄漏仅边际改善,因此DP对实际MIA防御的影响似乎有限。我们的发现揭示了实现更优效用-隐私权衡的可能性,并认为基于实证的攻击特异性隐私估计在调优实际隐私中可发挥关键作用。