Generative foundation models like Stable Diffusion comprise a diverse spectrum of knowledge in computer vision with the potential for transfer learning, e.g., via generating data to train student models for downstream tasks. This could circumvent the necessity of collecting labeled real-world data, thereby presenting a form of data-free knowledge distillation. However, the resultant student models show a significant drop in accuracy compared to models trained on real data. We investigate possible causes for this drop and focus on the role of the different layers of the student model. By training these layers using either real or synthetic data, we reveal that the drop mainly stems from the model's final layers. Further, we briefly investigate other factors, such as differences in data-normalization between synthetic and real, the impact of data augmentations, texture vs.\ shape learning, and assuming oracle prompts. While we find that some of those factors can have an impact, they are not sufficient to close the gap towards real data. Building upon our insights that mainly later layers are responsible for the drop, we investigate the data-efficiency of fine-tuning a synthetically trained model with real data applied to only those last layers. Our results suggest an improved trade-off between the amount of real training data used and the model's accuracy. Our findings contribute to the understanding of the gap between synthetic and real data and indicate solutions to mitigate the scarcity of labeled real data.
翻译:生成式基础模型(如Stable Diffusion)包含计算机视觉中多样化的知识谱系,具有迁移学习的潜力,例如通过生成数据来训练下游任务的学生模型。这可以避免收集标注真实数据的必要性,从而形成一种无数据知识蒸馏形式。然而,与基于真实数据训练的模型相比,此类学生模型的准确率显著下降。我们探究了这一下降的可能原因,重点关注学生模型不同层次的作用。通过使用真实数据或合成数据训练这些层次,我们发现下降主要源于模型的最终层。此外,我们简要研究了其他因素,如合成数据与真实数据之间数据归一化的差异、数据增强的影响、纹理与形状学习,以及假设最优提示。尽管我们发现某些因素可能产生影响,但不足以弥合与真实数据的差距。基于对主要下降源于后期层的洞察,我们研究了仅对最后几层应用真实数据微调合成训练模型的数据效率。结果表明,在使用的真实训练数据量与模型准确率之间存在更优的权衡。我们的发现有助于理解合成数据与真实数据之间的差距,并提出了缓解标注真实数据稀缺问题的解决方案。