Publicly available diabetic retinopathy (DR) datasets are imbalanced, containing limited numbers of images with DR. This imbalance contributes to overfitting when training machine learning classifiers. The impact of this imbalance is exacerbated as the severity of the DR stage increases, affecting the classifiers' diagnostic capacity. The imbalance can be addressed using Generative Adversarial Networks (GANs) to augment the datasets with synthetic images. Generating synthetic images is advantageous if high-quality and diversified images are produced. To evaluate the quality and diversity of synthetic images, several evaluation metrics, such as Multi-Scale Structural Similarity Index (MS-SSIM), Cosine Distance (CD), and Fr\'echet Inception Distance (FID) are used. Understanding the effectiveness of each metric in evaluating the quality and diversity of GAN-based synthetic images is critical to select images for augmentation. To date, there has been limited analysis of the appropriateness of these metrics in the context of biomedical imagery. This work contributes an empirical assessment of these evaluation metrics as applied to synthetic Proliferative DR imagery generated by a Deep Convolutional GAN (DCGAN). Furthermore, the metrics' capacity to indicate the quality and diversity of synthetic images and a correlation with classifier performance is undertaken. This enables a quantitative selection of synthetic imagery and an informed augmentation strategy. Results indicate that FID is suitable for evaluating the quality, while MS-SSIM and CD are suitable for evaluating the diversity of synthetic imagery. Furthermore, the superior performance of Convolutional Neural Network (CNN) and EfficientNet classifiers, as indicated by the F1 and AUC scores, for the augmented datasets demonstrates the efficacy of synthetic imagery to augment the imbalanced dataset.
翻译:公开可用的糖尿病视网膜病变(DR)数据集存在类别不平衡问题,其中包含的DR图像数量有限。这种不平衡现象在训练机器学习分类器时会导致过拟合问题。随着DR病变严重程度的增加,不平衡的影响进一步加剧,从而影响分类器的诊断能力。利用生成对抗网络(GAN)生成合成图像来扩充数据集,可以有效解决这种不平衡问题。生成合成图像的优势在于能够产生高质量且多样化的图像。为了评估合成图像的质量和多样性,通常采用多尺度结构相似性指数(MS-SSIM)、余弦距离(CD)和弗雷歇初始距离(FID)等评价指标。理解各指标在评估基于GAN的合成图像质量与多样性方面的有效性,对于选择用于数据增强的图像至关重要。迄今为止,在生物医学图像领域,对这些指标适用性的分析仍十分有限。本研究对深度卷积生成对抗网络(DCGAN)生成的增殖性DR合成图像进行了实证评估。此外,还探讨了这些指标表征合成图像质量与多样性的能力,及其与分类器性能的相关性。这实现了对合成图像的定量筛选和基于数据增强策略的理性决策。结果表明,FID适用于评估图像质量,而MS-SSIM和CD适用于评估图像多样性。此外,基于扩充数据集的卷积神经网络(CNN)和EfficientNet分类器在F1分数和AUC分数上表现更优,验证了合成图像在增强不平衡数据集方面的有效性。