Diffusion models excel at generating photorealistic images from text-queries. Naturally, many approaches have been proposed to use these generative abilities to augment training datasets for downstream tasks, such as classification. However, diffusion models are themselves trained on large noisily supervised, but nonetheless, annotated datasets. It is an open question whether the generalization capabilities of diffusion models beyond using the additional data of the pre-training process for augmentation lead to improved downstream performance. We perform a systematic evaluation of existing methods to generate images from diffusion models and study new extensions to assess their benefit for data augmentation. While we find that personalizing diffusion models towards the target data outperforms simpler prompting strategies, we also show that using the training data of the diffusion model alone, via a simple nearest neighbor retrieval procedure, leads to even stronger downstream performance. Overall, our study probes the limitations of diffusion models for data augmentation but also highlights its potential in generating new training data to improve performance on simple downstream vision tasks.
翻译:扩散模型在根据文本查询生成逼真图像方面表现卓越。自然,许多方法被提出利用这些生成能力来增强下游任务(如分类)的训练数据集。然而,扩散模型本身是在大规模带噪声监督但仍带有标注的数据集上训练的。一个开放的问题是:扩散模型在预训练过程中利用额外数据进行增强的泛化能力,是否能够带来下游性能的提升?我们对现有基于扩散模型生成图像的方法进行了系统评估,并研究了新的扩展方案以评估其对数据增强的益处。我们发现,将扩散模型针对目标数据进行个性化调优优于简单的提示策略,但同时也表明,仅通过简单的最近邻检索过程使用扩散模型的训练数据,就能实现更强的下游性能。总体而言,本研究探究了扩散模型在数据增强中的局限性,同时也凸显了其在生成新训练数据以改善简单下游视觉任务性能方面的潜力。