Verifying the robustness of machine learning models against evasion attacks at test time is an important research problem. Unfortunately, prior work established that this problem is NP-hard for decision tree ensembles, hence bound to be intractable for specific inputs. In this paper, we identify a restricted class of decision tree ensembles, called large-spread ensembles, which admit a security verification algorithm running in polynomial time. We then propose a new approach called verifiable learning, which advocates the training of such restricted model classes which are amenable for efficient verification. We show the benefits of this idea by designing a new training algorithm that automatically learns a large-spread decision tree ensemble from labelled data, thus enabling its security verification in polynomial time. Experimental results on public datasets confirm that large-spread ensembles trained using our algorithm can be verified in a matter of seconds, using standard commercial hardware. Moreover, large-spread ensembles are more robust than traditional ensembles against evasion attacks, at the cost of an acceptable loss of accuracy in the non-adversarial setting.
翻译:测试时抵御规避攻击的机器学习模型鲁棒性验证是一个重要的研究问题。然而,先前的研究已证明,决策树集成模型的该问题属于NP难问题,因此对特定输入必然难以处理。本文定义了一类受限的决策树集成——大间距集成,该类模型可在多项式时间内完成安全验证。我们提出名为"可验证学习"的新范式,倡导训练此类易于高效验证的受限模型类别。通过设计新的训练算法,我们展示了该思想的优势:该算法能从标注数据中自动学习大间距决策树集成,从而在多项式时间内实现其安全验证。在公开数据集上的实验结果表明,使用标准商用硬件,经由该算法训练的大间距集成可在数秒内完成验证。此外,相较于传统集成方法,大间距集成在对抗非对抗场景中可接受的精度损失为代价,展现出更强的规避攻击鲁棒性。