Machine Learning as a Service (MLaaS) platforms have gained popularity due to their accessibility, cost-efficiency, scalability, and rapid development capabilities. However, recent research has highlighted the vulnerability of cloud-based models in MLaaS to model extraction attacks. In this paper, we introduce FDINET, a novel defense mechanism that leverages the feature distribution of deep neural network (DNN) models. Concretely, by analyzing the feature distribution from the adversary's queries, we reveal that the feature distribution of these queries deviates from that of the model's training set. Based on this key observation, we propose Feature Distortion Index (FDI), a metric designed to quantitatively measure the feature distribution deviation of received queries. The proposed FDINET utilizes FDI to train a binary detector and exploits FDI similarity to identify colluding adversaries from distributed extraction attacks. We conduct extensive experiments to evaluate FDINET against six state-of-the-art extraction attacks on four benchmark datasets and four popular model architectures. Empirical results demonstrate the following findings FDINET proves to be highly effective in detecting model extraction, achieving a 100% detection accuracy on DFME and DaST. FDINET is highly efficient, using just 50 queries to raise an extraction alarm with an average confidence of 96.08% for GTSRB. FDINET exhibits the capability to identify colluding adversaries with an accuracy exceeding 91%. Additionally, it demonstrates the ability to detect two types of adaptive attacks.
翻译:机器学习即服务(Machine Learning as a Service, MLaaS)平台因其易用性、成本效益、可扩展性和快速开发能力而广受欢迎。然而,近期研究表明,MLaaS中基于云的模型易遭受模型提取攻击。本文提出了一种新型防御机制FDINET,该机制利用深度神经网络(DNN)模型的特征分布特性。具体而言,通过分析攻击者查询的特征分布,我们发现这些查询的特征分布与模型训练集的特征分布存在偏离。基于这一关键观察,我们提出了特征失真指数(Feature Distortion Index, FDI),一种用于定量测量接收查询特征分布偏差的度量指标。所提出的FDINET利用FDI训练二值检测器,并通过FDI相似性识别分布式提取攻击中的合谋攻击者。我们在四个基准数据集和四种主流模型架构上进行了大量实验,以评估FDINET对抗六种最新提取攻击的性能。实验结果表明:FDINET在检测模型提取方面表现出极高的有效性,对DFME和DaST的检测准确率达到100%;FDINET效率极高,仅需50次查询即可对GTSRB数据集发起提取警报,平均置信度达96.08%;FDINET识别合谋攻击者的准确率超过91%;此外,它还能检测两种类型的自适应攻击。