Black-box query attacks, which rely only on the output of the victim model, have proven to be effective in attacking deep learning models. However, existing black-box query attacks show low performance in a novel scenario where only a few queries are allowed. To address this issue, we propose gradient aligned attacks (GAA), which use the gradient aligned losses (GAL) we designed on the surrogate model to estimate the accurate gradient to improve the attack performance on the victim model. Specifically, we propose a gradient aligned mechanism to ensure that the derivatives of the loss function with respect to the logit vector have the same weight coefficients between the surrogate and victim models. Using this mechanism, we transform the cross-entropy (CE) loss and margin loss into gradient aligned forms, i.e. the gradient aligned CE or margin losses. These losses not only improve the attack performance of our gradient aligned attacks in the novel scenario but also increase the query efficiency of existing black-box query attacks. Through theoretical and empirical analysis on the ImageNet database, we demonstrate that our gradient aligned mechanism is effective, and that our gradient aligned attacks can improve the attack performance in the novel scenario by 16.1\% and 31.3\% on the $l_2$ and $l_{\infty}$ norms of the box constraint, respectively, compared to four latest transferable prior-based query attacks. Additionally, the gradient aligned losses also significantly reduce the number of queries required in these transferable prior-based query attacks by a maximum factor of 2.9 times. Overall, our proposed gradient aligned attacks and losses show significant improvements in the attack performance and query efficiency of black-box query attacks, particularly in scenarios where only a few queries are allowed.
翻译:黑盒查询攻击仅依赖受害模型输出,已被证明能有效攻击深度学习模型。然而,在仅允许少量查询的新场景中,现有黑盒查询攻击的性能较低。为解决此问题,我们提出梯度对齐攻击(GAA),该方法利用在替代模型上设计的梯度对齐损失(GAL)来估计精确梯度,从而提升对受害模型的攻击性能。具体而言,我们提出一种梯度对齐机制,确保损失函数对逻辑向量导数的权重系数在替代模型与受害模型之间保持一致。通过该机制,我们将交叉熵(CE)损失和边际损失转化为梯度对齐形式,即梯度对齐CE损失或边际损失。这些损失不仅提升了梯度对齐攻击在新场景中的攻击性能,还提高了现有黑盒查询攻击的查询效率。通过在ImageNet数据库上的理论与实证分析,我们证明了梯度对齐机制的有效性,并且与四种最新的可迁移先验查询攻击相比,我们的梯度对齐攻击在盒约束的$l_2$和$l_{\infty}$范数下,分别将新场景中的攻击性能提升了16.1%和31.3%。此外,梯度对齐损失还将这些可迁移先验查询攻击所需的查询次数最多降低了2.9倍。总体而言,我们提出的梯度对齐攻击与损失显著提升了黑盒查询攻击的攻击性能与查询效率,尤其在仅允许少量查询的场景中表现突出。