Diffusion models trained on large datasets can synthesize photo-realistic images of remarkable quality and diversity. However, attributing these images back to the training data-that is, identifying specific training examples which caused an image to be generated-remains a challenge. In this paper, we propose a framework that: (i) provides a formal notion of data attribution in the context of diffusion models, and (ii) allows us to counterfactually validate such attributions. Then, we provide a method for computing these attributions efficiently. Finally, we apply our method to find (and evaluate) such attributions for denoising diffusion probabilistic models trained on CIFAR-10 and latent diffusion models trained on MS COCO. We provide code at https://github.com/MadryLab/journey-TRAK .
翻译:基于大规模数据集训练的扩散模型能够合成具有卓越质量和多样性的照片级逼真图像。然而,将这些图像归因于训练数据——即识别出导致图像生成的特定训练样本——仍然是一项挑战。在本文中,我们提出一个框架,该框架:(i)在扩散模型背景下提供数据归因的形式化定义,以及(ii)允许我们对这类归因进行反事实验证。随后,我们提供一种高效计算这些归因的方法。最后,我们将该方法应用于在CIFAR-10上训练的降噪扩散概率模型和在MS COCO上训练的潜在扩散模型,以查找(并评估)此类归因。我们在https://github.com/MadryLab/journey-TRAK 提供代码。