Diffusion models are a class of generative models that generate high-quality samples, but at present it is difficult to characterize how they depend upon their training data. This difficulty raises scientific and regulatory questions, and is a consequence of the complexity of diffusion models and their sampling process. To analyze this dependence, we introduce Ablation Based Counterfactuals (ABC), a method of performing counterfactual analysis that relies on model ablation rather than model retraining. In our approach, we train independent components of a model on different but overlapping splits of a training set. These components are then combined into a single model, from which the causal influence of any training sample can be removed by ablating a combination of model components. We demonstrate how we can construct a model like this using an ensemble of diffusion models. We then use this model to study the limits of training data attribution by enumerating full counterfactual landscapes, and show that single source attributability diminishes with increasing training data size. Finally, we demonstrate the existence of unattributable samples.
翻译:扩散模型是一类能够生成高质量样本的生成模型,但目前难以刻画其对训练数据的依赖关系。这种困难引发了科学和监管层面的问题,其根源在于扩散模型及其采样过程的复杂性。为分析这种依赖关系,我们提出了基于消融的反事实推理方法,该方法通过模型消融而非重新训练来实现反事实分析。在我们的方法中,我们将模型的独立组件在不同但重叠的训练集划分上进行训练,随后将这些组件组合成单一模型。通过消融特定模型组件的组合,即可消除任意训练样本的因果影响。我们展示了如何利用扩散模型集成来构建此类模型。随后,我们通过枚举完整的反事实场景来研究训练数据归因的局限性,并证明单源可归因性随训练数据规模增加而减弱。最后,我们论证了不可归因样本的存在性。