Model stealing aims at inferring a victim model's functionality at a fraction of the original training cost. While the goal is clear, in practice the model's architecture, weight dimension, and original training data can not be determined exactly, leading to mutual uncertainty during stealing. In this work, we explicitly tackle this uncertainty by generating multiple possible networks and combining their predictions to improve the quality of the stolen model. For this, we compare five popular uncertainty quantification models in a model stealing task. Surprisingly, our results indicate that the considered models only lead to marginal improvements in terms of label agreement (i.e., fidelity) to the stolen model. To find the cause of this, we inspect the diversity of the model's prediction by looking at the prediction variance as a function of training iterations. We realize that during training, the models tend to have similar predictions, indicating that the network diversity we wanted to leverage using uncertainty quantification models is not (high) enough for improvements on the model stealing task.
翻译:模型窃取旨在以原始训练成本的一小部分推断受害者模型的功能。尽管目标明确,但实际中无法精确确定模型架构、权重维度及原始训练数据,导致窃取过程存在相互不确定性。本研究通过生成多个可能网络并整合其预测结果,显式处理此类不确定性以提升窃取模型质量。为此,我们在模型窃取任务中比较了五种主流不确定性量化模型。令人惊讶的是,实验结果表明,这些模型在标签一致性(即保真度)上仅带来边际改善。为探究成因,我们通过观测预测方差随训练迭代的变化来审视模型预测的多样性。研究发现,训练过程中模型预测趋于相似,表明我们试图借助不确定性量化模型利用的网络多样性,在模型窃取任务中尚不足以产生显著改进。