Biometric authentication service providers often claim that it is not possible to reverse-engineer a user's raw biometric sample, such as a fingerprint or a face image, from its mathematical (feature-space) representation. In this paper, we investigate this claim on the specific example of deep neural network (DNN) embeddings. Inversion of DNN embeddings has been investigated for explaining deep image representations or synthesizing normalized images. Existing studies leverage full access to all layers of the original model, as well as all possible information on the original dataset. For the biometric authentication use case, we need to investigate this under adversarial settings where an attacker has access to a feature-space representation but no direct access to the exact original dataset nor the original learned model. Instead, we assume varying degree of attacker's background knowledge about the distribution of the dataset as well as the original learned model (architecture and training process). In these cases, we show that the attacker can exploit off-the-shelf DNN models and public datasets, to mimic the behaviour of the original learned model to varying degrees of success, based only on the obtained representation and attacker's prior knowledge. We propose a two-pronged attack that first infers the original DNN by exploiting the model footprint on the embedding, and then reconstructs the raw data by using the inferred model. We show the practicality of the attack on popular DNNs trained for two prominent biometric modalities, face and fingerprint recognition. The attack can effectively infer the original recognition model (mean accuracy 83\% for faces, 86\% for fingerprints), and can craft effective biometric reconstructions that are successfully authenticated with 1-vs-1 authentication accuracy of up to 92\% for some models.
翻译:生物特征认证服务提供商常声称,无法从用户的数学(特征空间)表示中反向工程出原始生物特征样本(如指纹或人脸图像)。本文以深度神经网络嵌入为具体案例,对该主张进行探究。深度神经网络嵌入的反演技术此前已被用于解释深度图像表征或合成归一化图像。现有研究需完全访问原始模型的所有层及原始数据集的所有可能信息。针对生物特征认证场景,我们需在对抗性设置下展开研究,即攻击者可访问特征空间表示,但无法直接获取精确的原始数据集或原始学习模型。取而代之,我们假定攻击者对数据集分布以及原始学习模型(架构与训练过程)具有不同程度的背景知识。在此类情况下,我们证明攻击者可利用现成的深度神经网络模型与公开数据集,仅凭所获取的特征表示与先验知识,以不同程度的成功率模仿原始学习模型的行为。我们提出双阶段攻击策略:首先通过嵌入层的模型足迹推断原始深度神经网络,再利用推断出的模型重建原始数据。我们以人脸与指纹识别两类主流生物特征模态训练的深度神经网络为例,展示了该攻击的实用性。该攻击能有效推断原始识别模型(人脸平均准确率83%,指纹平均准确率86%),并可生成有效的生物特征重建样本,部分模型的一对一认证成功率高达92%。