Large annotated datasets are required for training deep learning models, but in medical imaging data sharing is often complicated due to ethics, anonymization and data protection legislation (e.g. the general data protection regulation (GDPR)). Generative AI models, such as generative adversarial networks (GANs) and diffusion models, can today produce very realistic synthetic images, and can potentially facilitate data sharing as GDPR should not apply for medical images which do not belong to a specific person. However, in order to share synthetic images it must first be demonstrated that they can be used for training different networks with acceptable performance. Here, we therefore comprehensively evaluate four GANs (progressive GAN, StyleGAN 1-3) and a diffusion model for the task of brain tumor segmentation. Our results show that segmentation networks trained on synthetic images reach Dice scores that are 80\% - 90\% of Dice scores when training with real images, but that memorization of the training images can be a problem for diffusion models if the original dataset is too small. Furthermore, we demonstrate that common metrics for evaluating synthetic images, Fr\'echet inception distance (FID) and inception score (IS), do not correlate well with the obtained performance when using the synthetic images for training segmentation networks.
翻译:训练深度学习模型需要大量标注数据集,但在医学影像领域,数据共享常因伦理、匿名化及数据保护法规(如通用数据保护条例GDPR)而变得复杂。生成式AI模型,例如生成对抗网络(GANs)和扩散模型,如今能够生成高度逼真的合成图像,并有可能促进数据共享——因为GDPR不应适用于不属于特定个人的医学图像。然而,要共享合成图像,必须首先证明它们可用于训练不同网络且达到可接受的性能。为此,我们系统评估了四种GAN(渐进式GAN、StyleGAN 1-3)和一种扩散模型在脑肿瘤分割任务中的表现。结果表明,使用合成图像训练的分割网络,其Dice得分达到使用真实图像训练时的80%-90%,但若原始数据集过小,扩散模型可能存在对训练图像的记忆问题。此外,我们发现常用合成图像评价指标——弗雷歇初始距离(FID)和初始分数(IS)——与使用合成图像训练分割网络所获性能之间缺乏良好的相关性。