In this work, we highlight and perform a comprehensive study on calibration attacks, a form of adversarial attacks that aim to trap victim models to be heavily miscalibrated without altering their predicted labels, hence endangering the trustworthiness of the models and follow-up decision making based on their confidence. We propose four typical forms of calibration attacks: underconfidence, overconfidence, maximum miscalibration, and random confidence attacks, conducted in both the black-box and white-box setups. We demonstrate that the attacks are highly effective on both convolutional and attention-based models: with a small number of queries, they seriously skew confidence without changing the predictive performance. Given the potential danger, we further investigate the effectiveness of a wide range of adversarial defence and recalibration methods, including our proposed defences specifically designed for calibration attacks to mitigate the harm. From the ECE and KS scores, we observe that there are still significant limitations in handling calibration attacks. To the best of our knowledge, this is the first dedicated study that provides a comprehensive investigation on calibration-focused attacks. We hope this study helps attract more attention to these types of attacks and hence hamper their potential serious damages. To this end, this work also provides detailed analyses to understand the characteristics of the attacks.
翻译:本研究聚焦并系统性地探讨了校准攻击——一种旨在误导受害者模型严重失准但不改变其预测标签的对抗攻击形式,从而威胁模型的可信度及其基于置信度的后续决策机制。我们提出四种典型校准攻击形式:欠置信攻击、过置信攻击、最大失准攻击与随机置信攻击,并在黑盒与白盒两种设置下实施。实验表明,这类攻击对基于卷积和注意力机制的模型均高度有效:仅需少量查询即可严重扭曲模型置信度,同时保持预测性能不变。鉴于其潜在危害,我们进一步评估了多种对抗防御与再校准方法的有效性,包括针对校准攻击专门设计的防御策略。根据ECE与KS评分,我们发现现有方法在处理校准攻击时仍存在显著局限性。据我们所知,这是首个专门针对置信度导向攻击开展全面研究的系统性工作。我们期待本研究能引发学界对此类攻击的更多关注,从而遏制其可能造成的严重危害。为此,本文还提供了详尽的特征分析以深入理解攻击机理。