Recent research has shown it is possible to perform zero-shot classification tasks by training a classifier with synthetic data generated by a diffusion model. However, the performance of this approach is still inferior to that of recent vision-language models. It has been suggested that the reason for this is a domain gap between the synthetic and real data. In our work, we show that this domain gap is not the main issue, and that diversity in the synthetic dataset is more important. We propose a $\textit{bag of tricks}$ to improve diversity and are able to achieve performance on par with one of the vision-language models, CLIP. More importantly, this insight allows us to endow zero-shot classification capabilities on any classification model.
翻译:近期研究表明,通过扩散模型生成的合成数据训练分类器可实现零样本分类任务。然而,该方法的性能仍逊于最新的视觉-语言模型。已有观点认为其原因在于合成数据与真实数据之间存在领域差异。本研究发现,领域差异并非主要问题,合成数据集的多样性更为关键。我们提出一套$\textit{技巧集}$以增强多样性,并成功实现与视觉-语言模型CLIP相当的性能。更重要的是,这一发现使我们能够为任意分类模型赋予零样本分类能力。