Diffusion-based models have shown the merits of generating high-quality visual data while preserving better diversity in recent studies. However, such observation is only justified with curated data distribution, where the data samples are nicely pre-processed to be uniformly distributed in terms of their labels. In practice, a long-tailed data distribution appears more common and how diffusion models perform on such class-imbalanced data remains unknown. In this work, we first investigate this problem and observe significant degradation in both diversity and fidelity when the diffusion model is trained on datasets with class-imbalanced distributions. Especially in tail classes, the generations largely lose diversity and we observe severe mode-collapse issues. To tackle this problem, we set from the hypothesis that the data distribution is not class-balanced, and propose Class-Balancing Diffusion Models (CBDM) that are trained with a distribution adjustment regularizer as a solution. Experiments show that images generated by CBDM exhibit higher diversity and quality in both quantitative and qualitative ways. Our method benchmarked the generation results on CIFAR100/CIFAR100LT dataset and shows outstanding performance on the downstream recognition task.
翻译:基于扩散的模型在近年研究中展现出了生成高质量视觉数据同时保持较好多样性的优点。然而,这一观察仅在经过精心预处理使得数据样本标签均匀分布的数据分布下成立。在实践中,长尾数据分布更为常见,扩散模型在此类类别不平衡数据上的表现尚不明确。本文首先探究该问题,观察到当扩散模型在类别不平衡分布的数据集上训练时,其多样性和保真度均显著下降。特别是在尾部类别中,生成结果严重丧失多样性,并出现严重的模式崩溃问题。为解决该问题,我们从数据分布非平衡的假设出发,提出通过分布调整正则化训练得到类平衡扩散模型(CBDM)作为解决方案。实验表明,CBDM生成的图像在定量和定性两方面均展现出更高的多样性和质量。我们的方法在CIFAR100/CIFAR100LT数据集上建立了生成结果基准,并在下游识别任务上展现出卓越性能。