Generating medical images from human-drawn free-hand sketches holds promise for various important medical imaging applications. Due to the extreme difficulty in collecting free-hand sketch data in the medical domain, most deep learning-based methods have been proposed to generate medical images from the synthesized sketches (e.g., edge maps or contours of segmentation masks from real images). However, these models often fail to generalize on the free-hand sketches, leading to unsatisfactory results. In this paper, we propose a practical free-hand sketch-to-image generation model called Sketch2MedI that learns to represent sketches in StyleGAN's latent space and generate medical images from it. Thanks to the ability to encode sketches into this meaningful representation space, Sketch2MedI only requires synthesized sketches for training, enabling a cost-effective learning process. Our Sketch2MedI demonstrates a robust generalization to free-hand sketches, resulting in high-quality and realistic medical image generations. Comparative evaluations of Sketch2MedI against the pix2pix, CycleGAN, UNIT, and U-GAT-IT models show superior performance in generating pharyngeal images, both quantitative and qualitative across various metrics.
翻译:从手绘草图生成医学图像在多种重要医学成像应用中具有广阔前景。由于医学领域手绘草图数据收集极为困难,现有基于深度学习的方法大多采用合成草图(如真实图像的边缘图或分割掩膜轮廓)进行医学图像生成。然而,这些模型往往难以泛化至真实手绘草图,导致生成效果不佳。本文提出一种实用的手绘草图到图像生成模型Sketch2MedI,该模型学习在StyleGAN的潜在空间中表示草图并据此生成医学图像。得益于将草图编码至该语义表示空间的能力,Sketch2MedI仅需合成草图进行训练,实现了高效低成本的学习过程。实验表明Sketch2MedI对手绘草图具有鲁棒的泛化能力,能生成高质量且逼真的医学图像。通过与pix2pix、CycleGAN、UNIT及U-GAT-IT模型的对比评估,Sketch2MedI在咽部图像生成任务中,各项定量与定性指标均表现出优越性能。