With the wide-spread application of machine learning models, it has become critical to study the potential data leakage of models trained on sensitive data. Recently, various membership inference (MI) attacks are proposed to determine if a sample was part of the training set or not. The question is whether these attacks can be reliably used in practice. We show that MI models frequently misclassify neighboring nonmember samples of a member sample as members. In other words, they have a high false positive rate on the subpopulations of the exact member samples that they can identify. We then showcase a practical application of MI attacks where this issue has a real-world repercussion. Here, MI attacks are used by an external auditor (investigator) to show to a judge/jury that an auditee unlawfully used sensitive data. Due to the high false positive rate of MI attacks on member's subpopulations, auditee challenges the credibility of the auditor by revealing the performance of the MI attacks on these subpopulations. We argue that current membership inference attacks can identify memorized subpopulations, but they cannot reliably identify which exact sample in the subpopulation was used during the training.
翻译:随着机器学习模型的广泛应用,研究针对敏感数据训练的模型是否可能泄露数据已成为关键问题。近年来,多种成员推理攻击被提出,旨在判断某个样本是否属于训练集。问题在于,这些攻击在实践中能否被可靠利用。我们证明,成员推理模型经常将某个成员样本的邻近非成员样本错误分类为成员。换言之,对于它们能够识别的特定成员样本子群体,这些模型具有较高的假阳性率。随后,我们展示了一个实际应用场景,其中该问题会产生现实影响:外部审计员使用成员推理攻击向法官/陪审团证明被审计方非法使用了敏感数据。由于成员推理攻击在成员子群体上的高假阳性率,被审计方通过揭示这些子群体上的攻击性能来质疑审计员的可靠性。我们认为,当前的成员推理攻击能够识别被记忆的子群体,但无法可靠地判断子群体中哪个具体样本在训练过程中被使用。