As deep neural networks are increasingly deployed in sensitive application domains, such as healthcare and security, it's necessary to understand what kind of sensitive information can be inferred from these models. Existing model-targeted attacks all assume the attacker has known the application domain or training data distribution, which plays an essential role in successful attacks. Can removing the domain information from model APIs protect models from these attacks? This paper studies this critical problem. Unfortunately, even with minimal knowledge, i.e., accessing the model as an unnamed function without leaking the meaning of input and output, the proposed adaptive domain inference attack (ADI) can still successfully estimate relevant subsets of training data. We show that the extracted relevant data can significantly improve, for instance, the performance of model-inversion attacks. Specifically, the ADI method utilizes a concept hierarchy built on top of a large collection of available public and private datasets and a novel algorithm to adaptively tune the likelihood of leaf concepts showing up in the unseen training data. The ADI attack not only extracts partial training data at the concept level, but also converges fast and requires much fewer target-model accesses than another domain inference attack, GDI.
翻译:随着深度神经网络在医疗和安全等敏感应用领域中的部署日益增多,理解这些模型可能泄露何种敏感信息变得至关重要。现有的模型定向攻击均假设攻击者已知应用领域或训练数据分布,而这在成功实施攻击中起着关键作用。能否从模型应用程序接口(API)中移除域信息来保护模型免受这些攻击?本文研究了这一关键问题。不幸的是,即使攻击者掌握的信息极少,仅能将模型作为未命名函数访问而无法获知输入和输出的含义,所提出的自适应域推断攻击(ADI)仍能成功估计训练数据的相关子集。我们表明,提取的相关数据可显著提升例如模型逆向攻击的性能。具体而言,ADI方法利用基于大量公开和私有数据集构建的概念层次结构,并结合一种新颖算法自适应调整未见训练数据中叶子概念出现的似然度。ADI攻击不仅能在概念层面提取部分训练数据,而且收敛速度快,且相比另一种域推断攻击GDI,其所需的靶模型访问次数更少。