We introduce the $ARMOR_D$ methods as novel approaches to enhancing the adversarial robustness of deep learning models. These methods are based on a new class of optimal-transport-regularized divergences, constructed via an infimal convolution between an information divergence and an optimal-transport (OT) cost. We use these as tools to enhance adversarial robustness by maximizing the expected loss over a neighborhood of distributions, a technique known as distributionally robust optimization. Viewed as a tool for constructing adversarial samples, our method allows samples to be both transported, according to the OT cost, and re-weighted, according to the information divergence. We demonstrate the effectiveness of our method on malware detection and image recognition applications and find that, to our knowledge, it outperforms existing methods at enhancing the robustness against adversarial attacks. $ARMOR_D$ yields the robustified accuracy of $98.29\%$ against $FGSM$ and $98.18\%$ against $PGD^{40}$ on the MNIST dataset, reducing the error rate by more than $19.7\%$ and $37.2\%$ respectively compared to prior methods. Similarly, in malware detection, a discrete (binary) data domain, $ARMOR_D$ improves the robustified accuracy under $rFGSM^{50}$ attack compared to the previous best-performing adversarial training methods by $37.0\%$ while lowering false negative and false positive rates by $51.1\%$ and $57.53\%$, respectively.
翻译:我们提出$ARMOR_D$方法,作为增强深度学习模型对抗鲁棒性的新途径。这些方法基于一类新型最优传输正则化散度,通过信息散度与最优传输(OT)成本之间的下确界卷积构造而成。我们将其作为工具,通过最大化分布邻域内的期望损失来增强对抗鲁棒性,该技术被称为分布鲁棒优化。从构造对抗样本的角度来看,我们的方法允许样本既根据OT成本进行迁移,又根据信息散度进行重新加权。我们在恶意软件检测和图像识别应用中验证了该方法的有效性,并发现据我们所知,它在增强对抗攻击鲁棒性方面优于现有方法。在MNIST数据集上,$ARMOR_D$针对$FGSM$的鲁棒准确率达到$98.29\%$,针对$PGD^{40}$的鲁棒准确率达到$98.18\%$,与先前方法相比,错误率分别降低了$19.7\%$以上和$37.2\%$以上。类似地,在恶意软件检测这一离散(二值)数据领域中,$ARMOR_D$在$rFGSM^{50}$攻击下的鲁棒准确率相比先前表现最佳的对抗训练方法提升了$37.0\%$,同时将假阴性率和假阳性率分别降低了$51.1\%$和$57.53\%$。