Adversarial attacks are a potential threat to machine learning models, as they can cause the model to make incorrect predictions by introducing imperceptible perturbations to the input data. While extensively studied in unstructured data like images, their application to structured data like tabular data presents unique challenges due to the heterogeneity and intricate feature interdependencies of tabular data. Imperceptibility in tabular data involves preserving data integrity while potentially causing misclassification, underscoring the need for tailored imperceptibility criteria for tabular data. However, there is currently a lack of standardised metrics for assessing adversarial attacks specifically targeted at tabular data. To address this gap, we derive a set of properties for evaluating the imperceptibility of adversarial attacks on tabular data. These properties are defined to capture seven perspectives of perturbed data: proximity to original inputs, sparsity of alterations, deviation to datapoints in the original dataset, sensitivity of altering sensitive features, immutability of perturbation, feasibility of perturbed values and intricate feature interdepencies among tabular features. Furthermore, we conduct both quantitative empirical evaluation and case-based qualitative examples analysis for seven properties. The evaluation reveals a trade-off between attack success and imperceptibility, particularly concerning proximity, sensitivity, and deviation. Although no evaluated attacks can achieve optimal effectiveness and imperceptibility simultaneously, unbounded attacks prove to be more promised for tabular data in crafting imperceptible adversarial examples. The study also highlights the limitation of evaluated algorithms in controlling sparsity effectively. We suggest incorporating a sparsity metric in future attack design to regulate the number of perturbed features.
翻译:对抗攻击是机器学习模型的潜在威胁,其通过对输入数据施加难以察觉的扰动,可导致模型做出错误预测。尽管在图像等非结构化数据中已得到广泛研究,但针对表格数据这类结构化数据的对抗攻击,由于表格数据特征的异构性及复杂的特征间相互依赖关系,呈现出独特的挑战。表格数据中的不可感知性要求在可能导致错误分类的同时保持数据完整性,这凸显了为表格数据定制不可感知性标准的必要性。然而,目前尚缺乏专门用于评估针对表格数据的对抗攻击的标准化度量指标。为填补这一空白,我们推导出一组用于评估表格数据对抗攻击不可感知性的属性。这些属性旨在从七个维度刻画扰动后的数据:与原始输入的接近程度、修改的稀疏性、相对于原始数据集中数据点的偏离程度、修改敏感特征的敏感性、扰动的不可变性、扰动值的可行性以及表格特征间复杂的特征相互依赖性。此外,我们针对这七个属性开展了定量的实证评估和基于案例的定性示例分析。评估揭示了攻击成功率和不可感知性之间的权衡关系,尤其是在接近性、敏感性和偏离性方面。尽管所评估的攻击方法均无法同时实现最优的攻击效果与不可感知性,但无界攻击在构建难以察觉的对抗样本方面对表格数据展现出更大潜力。研究还揭示了所评估算法在有效控制稀疏性方面的局限性。我们建议在未来的攻击设计中引入稀疏性度量指标,以调控被修改特征的数量。