Data poisoning considers cases when an adversary maliciously inserts and removes training data to manipulate the behavior of machine learning algorithms. Traditional threat models of data poisoning center around a single metric, the number of poisoned samples. In consequence, existing defenses are essentially vulnerable in practice when poisoning more samples remains a feasible option for attackers. To address this issue, we leverage timestamps denoting the birth dates of data, which are often available but neglected in the past. Benefiting from these timestamps, we propose a temporal threat model of data poisoning and derive two novel metrics, earliness and duration, which respectively measure how long an attack started in advance and how long an attack lasted. With these metrics, we define the notions of temporal robustness against data poisoning, providing a meaningful sense of protection even with unbounded amounts of poisoned samples. We present a benchmark with an evaluation protocol simulating continuous data collection and periodic deployments of updated models, thus enabling empirical evaluation of temporal robustness. Lastly, we develop and also empirically verify a baseline defense, namely temporal aggregation, offering provable temporal robustness and highlighting the potential of our temporal modeling of data poisoning.
翻译:数据投毒考虑的是恶意攻击者通过插入和移除训练数据来操控机器学习算法行为的情形。传统数据投毒的威胁模型主要围绕单一指标——投毒样本数量展开。因此,当攻击者选择投毒更多样本仍属可行选项时,现有防御机制在实战中根本不堪一击。为解决此问题,我们利用数据生成时间戳这一常被忽视但广泛可用的信息。借助时间戳,我们提出数据投毒的时序威胁模型,并推导出两个新指标:提前量与持续时长,分别度量攻击提前发起的时间长度和攻击持续的时间跨度。基于这些指标,我们定义了数据投毒的时序鲁棒性概念,即便面对无数量限制的投毒样本,仍能提供有意义的防护。我们构建了一个基准测试框架,配备模拟连续数据收集和定期部署更新模型的评估协议,从而实现对时序鲁棒性的实证评估。最后,我们开发并实验验证了基线防御方法——时序聚合,该方法能提供可证明的时序鲁棒性,凸显了数据投毒时序建模的潜力。