Decision-based iterative fragile watermarking for model integrity verification

Typically, foundation models are hosted on cloud servers to meet the high demand for their services. However, this exposes them to security risks, as attackers can modify them after uploading to the cloud or transferring from a local system. To address this issue, we propose an iterative decision-based fragile watermarking algorithm that transforms normal training samples into fragile samples that are sensitive to model changes. We then compare the output of sensitive samples from the original model to that of the compromised model during validation to assess the model's completeness.The proposed fragile watermarking algorithm is an optimization problem that aims to minimize the variance of the predicted probability distribution outputed by the target model when fed with the converted sample.We convert normal samples to fragile samples through multiple iterations. Our method has some advantages: (1) the iterative update of samples is done in a decision-based black-box manner, relying solely on the predicted probability distribution of the target model, which reduces the risk of exposure to adversarial attacks, (2) the small-amplitude multiple iterations approach allows the fragile samples to perform well visually, with a PSNR of 55 dB in TinyImageNet compared to the original samples, (3) even with changes in the overall parameters of the model of magnitude 1e-4, the fragile samples can detect such changes, and (4) the method is independent of the specific model structure and dataset. We demonstrate the effectiveness of our method on multiple models and datasets, and show that it outperforms the current state-of-the-art.

翻译：通常，基础模型托管在云服务器上以满足对其服务的高需求。然而，这使它们面临安全风险，因为攻击者可以在将模型上传到云端或从本地系统转移后对其进行修改。为解决这一问题，我们提出了一种迭代的基于决策的脆弱水印算法，该算法将正常训练样本转换为对模型变化敏感的脆弱样本。随后，我们在验证过程中比较原始模型与受损模型对敏感样本的输出，以评估模型的完整性。所提出的脆弱水印算法是一个优化问题，旨在最小化目标模型在输入转换后的样本时输出的预测概率分布的方差。我们通过多次迭代将正常样本转换为脆弱样本。该方法具有以下优点：(1) 样本的迭代更新以基于决策的黑盒方式进行，仅依赖于目标模型的预测概率分布，从而降低了暴露于对抗性攻击的风险；(2) 小幅度的多次迭代方法使脆弱样本在视觉上表现良好，在TinyImageNet上与原样本相比，峰值信噪比（PSNR）达到55 dB；(3) 即使模型整体参数变化幅度为1e-4，脆弱样本也能检测到此类变化；(4) 该方法独立于具体的模型结构和数据集。我们在多个模型和数据集上展示了该方法的有效性，并表明其性能优于当前最先进的方法。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

【硬核课】机器人学习课程，UT Austin朱玉可博士讲述自主机器人的人工智能与机器学习机器学习算法

专知会员服务

41+阅读 · 2020年9月21日

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日