Hand-crafted image quality metrics, such as PSNR and SSIM, are commonly used to evaluate model privacy risk under reconstruction attacks. Under these metrics, reconstructed images that are determined to resemble the original one generally indicate more privacy leakage. Images determined as overall dissimilar, on the other hand, indicate higher robustness against attack. However, there is no guarantee that these metrics well reflect human opinions, which, as a judgement for model privacy leakage, are more trustworthy. In this paper, we comprehensively study the faithfulness of these hand-crafted metrics to human perception of privacy information from the reconstructed images. On 5 datasets ranging from natural images, faces, to fine-grained classes, we use 4 existing attack methods to reconstruct images from many different classification models and, for each reconstructed image, we ask multiple human annotators to assess whether this image is recognizable. Our studies reveal that the hand-crafted metrics only have a weak correlation with the human evaluation of privacy leakage and that even these metrics themselves often contradict each other. These observations suggest risks of current metrics in the community. To address this potential risk, we propose a learning-based measure called SemSim to evaluate the Semantic Similarity between the original and reconstructed images. SemSim is trained with a standard triplet loss, using an original image as an anchor, one of its recognizable reconstructed images as a positive sample, and an unrecognizable one as a negative. By training on human annotations, SemSim exhibits a greater reflection of privacy leakage on the semantic level. We show that SemSim has a significantly higher correlation with human judgment compared with existing metrics. Moreover, this strong correlation generalizes to unseen datasets, models and attack methods.
翻译:人工设计的图像质量指标(如PSNR和SSIM)常用于评估模型在重建攻击下的隐私风险。在这些指标下,与原始图像相似的重建图像通常意味着更高的隐私泄露,而整体不相似的图像则表明对攻击具有更强的鲁棒性。然而,这些指标能否充分反映人类意见——作为衡量模型隐私泄露的更可靠判断——仍缺乏保证。本文全面研究了这些人工设计指标对重建图像中隐私信息的人类感知的忠实程度。我们在5个数据集(涵盖自然图像、人脸和细粒度类别)上,使用4种现有攻击方法从多种不同分类模型中重建图像,并对每张重建图像邀请多位人类标注者评估其是否可被识别。研究表明,人工设计指标与人类对隐私泄露的评估仅存在弱相关性,且这些指标本身往往相互矛盾。这些发现揭示了当前社区常用指标的潜在风险。针对这一风险,我们提出一种基于学习的度量方法SemSim,用于评估原始图像与重建图像之间的语义相似性。该度量采用标准三元组损失训练,以原始图像为锚点,其可识别的重建图像为正样本,不可识别的重建图像为负样本。通过基于人类标注的训练,SemSim在语义层面更充分地反映了隐私泄露。实验表明,相比现有指标,SemSim与人类判断的相关性显著更高,且这一强相关性可泛化至未见数据集、模型和攻击方法。