The usage of medical image data for the training of large-scale machine learning approaches is particularly challenging due to its scarce availability and the costly generation of data annotations, typically requiring the engagement of medical professionals. The rapid development of generative models allows towards tackling this problem by leveraging large amounts of realistic synthetically generated data for the training process. However, randomly choosing synthetic samples, might not be an optimal strategy. In this work, we investigate the targeted generation of synthetic training data, in order to improve the accuracy and robustness of image classification. Therefore, our approach aims to guide the generative model to synthesize data with high epistemic uncertainty, since large measures of epistemic uncertainty indicate underrepresented data points in the training set. During the image generation we feed images reconstructed by an auto encoder into the classifier and compute the mutual information over the class-probability distribution as a measure for uncertainty.We alter the feature space of the autoencoder through an optimization process with the objective of maximizing the classifier uncertainty on the decoded image. By training on such data we improve the performance and robustness against test time data augmentations and adversarial attacks on several classifications tasks.
翻译:利用医学图像数据训练大规模机器学习方法面临特殊挑战,主要源于数据稀缺性及标注生成成本高昂——通常需要医疗专业人员参与。生成模型的快速发展为应对此问题提供了可能,即通过利用大量逼真的合成生成数据来辅助训练过程。然而,随机选择合成样本可能并非最优策略。本研究探索定向生成合成训练数据的方法,旨在提升图像分类的准确性与鲁棒性。为此,我们提出引导生成模型合成具有高认知不确定性的数据,因为较大的认知不确定性度量表明训练集中存在代表性不足的数据点。在图像生成过程中,我们将自编码器重建的图像输入分类器,并计算类别概率分布上的互信息作为不确定性度量。通过优化过程改变自编码器的特征空间,其目标在于最大化分类器对解码图像的不确定性。基于此类数据的训练,我们在多项分类任务中提升了模型性能,并增强了针对测试时数据增强与对抗攻击的鲁棒性。